The first time a customer exports a year of data, the request passes the host’s limit, 60 seconds on a Netlify synchronous function, and the export dies. How to run long tasks in the background comes down to four parts: the request records a job, a worker does the work, a job row holds the status, and the user hears when it is done.
How to run long tasks in the background: the four parts every version has
Long tasks run in the background of a web app through four parts: the request writes a job and answers at once with a job id, a worker outside the request does the work, the job row carries status and progress, and the user gets the result by polling, subscription or email. The browser only watches.
This is one of the controls in hardening SaaS applications for resilience, and it applies to three kinds of long work in particular: AI generation, exports and bulk email. My working rule for what counts as long: anything that can run past a few seconds on a slow day, or anything that has to survive the user leaving the page.
The table is my frame for the four parts. A version built on one Postgres table and a version built on a hosted queue both have all four.
| Part | What it does | The simplest version | What it grows into |
|---|---|---|---|
| The enqueue | The request validates the input, writes one job row or queue message, and answers straight away with the job id | An insert into a jobs table | A message queue that the worker reads from |
| The worker | A process outside the request claims jobs and does the work | A small always-on service, or a scheduled function that drains the table | Several workers sharing the table, or a job library |
| The status | Queued, running, succeeded or failed, with progress and an error message the user can read | Columns on the same job row | A status endpoint the page reads, with a progress bar |
| The delivery | The page polls or subscribes, and an email or notification goes out for anything the user will not wait for | The page asks for the job every few seconds | A realtime subscription plus a notification |
If the slow thing is an ordinary page load or API call that has simply become slow, it should get faster rather than move off the request, and that is a different job: API response time for an AI-built app.
What goes wrong without it
Lengthy work can exceed request limits or be interrupted when a browser disconnects. To the user both look the same, a spinner and then an error or nothing, but the causes differ, and so does the first thing to check.
The request times out on a long export
A long export that times out has hit the shortest clock between browser and code. How long a Lambda can run is capped at 15 minutes per invocation, but a REST API Gateway in front allows up to 29 seconds unless the quota is raised on a Regional or private API. Make the export a job, not a bigger limit.
Every layer between the browser and your code has a clock, and the one that runs out first decides. The model call works the same way, and the table for Netlify, Supabase Edge Functions and the OpenAI client sits in why an OpenAI API call times out. Vercel’s plan-by-plan ladder, FUNCTION_INVOCATION_TIMEOUT and maxDuration are in Vercel function timeout limits, causes and fixes. The clocks below cover three more places a request can run: AWS, Cloudflare and a server of your own behind nginx.
| Where the request runs | The clock | What happens when it passes | Source, checked |
|---|---|---|---|
| AWS Lambda | Function timeout: default 3 seconds, maximum “900 seconds (15 minutes)”; Lambda Managed Instances allow up to 90 minutes for asynchronous and event source mapping invocations (except Amazon MQ and Amazon DocumentDB) | The function times out | AWS Lambda quotas and timeout pages, 2026-10-04 |
| Amazon API Gateway, REST API | Integration timeout “50 milliseconds - 29 seconds for all integration types”; can be raised for Regional and private APIs, not for edge-optimized APIs | The gateway returns INTEGRATION_TIMEOUT, a 504 | API Gateway REST quotas and gateway response types, 2026-10-04 |
| Amazon API Gateway, HTTP API | Maximum integration timeout 30 seconds, cannot be increased | Not stated on AWS’s HTTP API quotas page | API Gateway HTTP API quotas, 2026-10-04 |
| Cloudflare Workers, HTTP request | No hard limit on duration while the client stays connected; CPU time 10 ms per request on Free, up to 5 minutes on Paid (default 30 seconds) | Past the CPU limit, Error 1102 with “Worker exceeded resource limits”; on disconnect, tasks “may be canceled” | Cloudflare Workers limits, 2026-10-04 |
| A server behind nginx | proxy_read_timeout, default 60s, timed between two successive reads, not over the whole response | ”the connection is closed” | nginx proxy module docs, 2026-10-04 |
Three details in those sources change what you do. AWS Lambda’s quotas give the 15-minute ceiling, but a function you never configured runs on the 3-second default, and API Gateway’s REST API quotas add that raising the integration timeout past 29 seconds “might require a reduction in your Region-level throttle quota for your account.” Cloudflare Workers’ limits set no hard limit on an HTTP request’s duration and cap its CPU time, and “Waiting on network requests (such as fetch() calls, KV reads, or database queries) does not count toward CPU time”, so an export that mostly waits on the database is limited in practice by how long the client stays connected (my reading). And nginx’s proxy_read_timeout resets on every read, so an export that streams rows as it goes can outlast 60 seconds, while one that builds the whole file before sending a byte cannot (my reading of the definition).
Export time grows with the data, and one query per row or a whole file built in memory makes a long export longer, but where the export runs matters more than how fast it is: make it a job first, then make it fast (my reading). Raising the limit is the wrong first move, because the next customer’s data can be bigger than this one’s.
The task stops when the user closes the browser
A task stops when the user closes the browser for one of three reasons: the page was driving the work, the work lived inside a request the platform may end when the client leaves, or it was pushed past the response on a timer it could not outlive. In each case the server never owned the task.
The three are my reading; where a vendor documents the mechanism, its own words follow. In the first, the page itself is the worker: a loop in the browser calls the API once per row or once per email, and when the tab closes the loop stops at whatever item it had reached. In the second, the work runs inside the request, and of HTTP-triggered Workers Cloudflare says: “When the client disconnects or the response is complete, tasks associated with that request may be canceled.”
The third looks like a fix and is not one. The work is handed to waitUntil so the response can go out first, but Vercel’s waitUntil reference says: “Promises passed to waitUntil() will have the same timeout as the function itself. If the function times out, the promises will be cancelled.” Cloudflare sets its own cap: waitUntil() “can extend execution for up to 30 seconds after the response is sent or the client disconnects.”
Picture a founder’s Vercel route that, to make a slow AI generation feel instant, answers the request at once and hands the generation to waitUntil, with nothing writing a job record. For the longest prompts the generation runs past the function’s timeout, the promise is canceled with the function, and the page has nothing to show. The lesson I take from it: waitUntil suits short tails such as logging, and long work needs a job row and a worker that outlive the function.
Warning the user with a beforeunload prompt does not solve any of this; MDN on beforeunload says the event “is not reliably fired, especially on mobile platforms.” The fix is ownership: from the moment the job row exists the server owns the task, and the browser only watches it.
How to do it on the common stacks
Four pieces follow: the jobs table, where the worker runs on your host, the AI generation case, and the tools you can adopt instead of writing your own.
The jobs table, the claim query and the status the user sees
A jobs table needs a status, a payload, progress, attempts, an error, the owner and tenant, timestamps and a lease expiry. Workers claim rows with FOR UPDATE SKIP LOCKED, which PostgreSQL offers for queue-like tables, and an expired lease returns a dead worker’s job to the queue.
The schema is mine; rename the columns to fit your app. On Supabase it goes in like any other table, with row level security on and no policy that lets a browser write to it directly.
create table jobs (
id bigint generated always as identity primary key,
type text not null,
payload jsonb not null,
status text not null default 'queued', -- queued, running, succeeded, failed
progress integer not null default 0,
attempts integer not null default 0,
result_path text,
error_message text,
created_by uuid not null, tenant_id uuid not null,
created_at timestamptz not null default now(),
started_at timestamptz, finished_at timestamptz,
lease_expires_at timestamptz
);
The claim query lets several workers share the table without two of them taking the same job. It runs as one statement, so the lock and the update happen together; $1 is the lease length, which should be longer than a job’s normal run.
update jobs
set status = 'running', attempts = attempts + 1,
started_at = now(), lease_expires_at = now() + $1::interval
where id = (select id from jobs where status = 'queued'
order by created_at limit 1
for update skip locked)
returning *;
PostgreSQL’s locking clause documentation is frank about the trade: “Skipping locked rows provides an inconsistent view of the data, so this is not suitable for general purpose work, but can be used to avoid lock contention with multiple consumers accessing a queue-like table.” A jobs table is that kind of table; reports and status reads should query it without SKIP LOCKED.
The lease is what keeps a crashed worker from leaving a job marked running forever. A small sweep, update jobs set status = 'queued' where status = 'running' and lease_expires_at < now(), puts expired jobs back in line, and a worker on a long job pushes its own lease_expires_at forward each time it writes progress, so a slow job is never mistaken for a dead one.
The status endpoint returns one job by id and only if created_by and tenant_id match the caller, so one customer can never read another’s job. The page polls it every few seconds or subscribes to changes on the row. An export’s result is a private file behind a signed link that expires, not an email attachment. And because a job can run twice after a crash, every effect it has must be safe to repeat, which is what an idempotency key for safe retries is for, along with retry limits and the record a failed job leaves.
Where the worker runs, by host
Each host offers a different background primitive, and each has its own clock. The primitive and clock columns come from each host’s own docs, read on the date in the last column; the “What you add” column is my advice.
| Host | The background primitive it offers | Its clock | What you add | Source, checked |
|---|---|---|---|---|
| Netlify | Background functions: the caller gets a 202 at once and the function keeps running; on Credit-based plans (Free, Personal and Pro) and on Enterprise plans | 15 minutes | The jobs table, so the page has a status to read | Netlify background functions and Functions configuration, 2026-10-04 |
| Vercel | waitUntil() from @vercel/functions; on Next.js 15.1 or above, after() from next/server | waitUntil(): the same timeout as the function itself; after(): the route’s “default or configured max duration” | A worker elsewhere for anything that can outlast the function | Vercel Functions package reference and Next.js after() reference, 2026-10-04 |
| Supabase | Queues on the pgmq extension; Cron on pg_cron; Edge Function background tasks with EdgeRuntime.waitUntil | Cron: “Each Job should run no more than 10 minutes”; Edge Functions: the function’s own limits, tabulated in the OpenAI timeout article | A worker that reads the queue, and the status columns | Supabase Queues, Cron and Background Tasks docs, 2026-10-04 |
| AWS | A Lambda function processing messages from an Amazon SQS queue | 15 minutes per invocation; up to 90 minutes on Lambda Managed Instances | A status table the page can read | AWS Lambda quotas and Using Lambda with Amazon SQS, 2026-10-04 |
| Cloudflare | Queues, with a Worker as the consumer | 15 minutes wall time per consumer invocation | The status store the page reads | Cloudflare Workers limits, 2026-10-04 |
| Render | A background worker: a service that runs continuously and takes no incoming network traffic | Not stated in Render’s background worker docs | A paid instance: Render’s free instances cover web services, Render Postgres and Render Key Value | Render background workers and free instances, 2026-10-04 |
| Any host | A scheduled function that drains the jobs table every few minutes | The host’s scheduled-function limit (30 seconds on Netlify) | A drain that claims a few jobs per run | Netlify Functions configuration, 2026-10-04 |
Some rows need more than a cell holds. Supabase Queues is “a Postgres-native durable Message Queue system with guaranteed delivery built on the pgmq database extension”, and Supabase Cron, built on pg_cron, can run jobs “anywhere from every second to once a year”, while EdgeRuntime.waitUntil keeps the function instance running “until the promise provided to waitUntil completes.” Render’s background workers “don’t receive any incoming network traffic” and instead “usually poll a task queue”, which is the jobs table above or a Key Value instance. On Vercel, after() lets you “schedule work that runs after the response has been sent”; for durable work, “Fix 3: move the long job off the request path entirely” in the Vercel timeout article is the next step. The scheduled drain is the plainest version of all, and a cron job every few minutes is enough for low volumes.
Hosted job services are a third route: they run the queue, the retries and the worker scheduling for you. Inngest’s docs describe it as “the platform for building durable workflows and agents that survive failures, without managing orchestration infrastructure”, and Trigger.dev’s docs call it “the open source platform for building and running durable AI agents and workflows in TypeScript.” I am not ranking either; each is another service to set up, weighed against a table you already have.
My working rule for choosing: under a few minutes per job and at low volume, the platform’s own primitive plus the jobs table is enough; once a job can outlast the platform’s clock, run an always-on worker.
Move AI generation to a background worker
AI generation moves to a background worker in six steps: store the inputs and return a job id, call the model with a deadline and bounded retries, record progress, save the output and token usage, fail with a readable reason, and notify the user. The browser shows the job’s status the whole time.
- 01 The request stores the prompt inputs on a new job row and returns the job id. Nothing calls the model yet.
- 02 The worker claims the job and calls the model with its own deadline, retrying rate limits and transient errors a bounded number of times with backoff.
- 03 For a generation with several steps, the worker writes progress to the row after each one, so the page can show where it is.
- 04 The worker stores the output, or its file location, and the token usage on the job row.
- 05 If the provider is down or keeps failing, the worker marks the job failed with a reason a person can read, and the page shows that reason.
- 06 The user is notified: the open page updates, and an email or in-app notice goes out for anything they did not wait for.
For step 2, the deadline is the same idea as how to set a timeout on fetch, and the waits between attempts follow what is exponential backoff. Count the client library’s own retries too. OpenAI’s Node SDK says “Certain errors will be automatically retried 2 times by default, with a short exponential backoff” and that “requests which time out will be retried twice by default”, and its maxRetries option configures or disables that. A worker that retries on top of that client multiplies the two counts, so set one of them to zero (my reading).
The provider-side choices, OpenAI’s background mode, streaming and the deadline to pick, are covered in the OpenAI timeout article linked earlier. For large offline batches that nobody is waiting on, OpenAI’s Batch API guide lists a “50% cost discount compared to synchronous APIs” and says “Each batch completes within 24 hours (and often more quickly).” Bulk email follows the same six steps, with the email provider’s send rate as the throttle on how fast the worker drains the queue (my reading).
Workflow orchestration tools, open-source job libraries, and when a table is enough
Open-source options fall into two groups. Job libraries such as BullMQ, pg-boss, Graphile Worker, Celery and Sidekiq run on Redis, Postgres or a broker you already run. Workflow orchestrators such as Temporal and Apache Airflow are separate systems for multi-step, long-lived processes; a small SaaS usually starts with the first.
Each row below is from the project’s own repository, with the license from its license file.
| Tool | What it runs on | Language | What it says it is for | License |
|---|---|---|---|---|
| BullMQ | Redis or PostgreSQL | TypeScript; for Node.js, Python, .NET, Elixir, Rust and PHP | ”Message Queue and Batch processing” | MIT |
| pg-boss | Postgres | TypeScript, for Node.js | ”Queueing jobs in Postgres from Node.js” | MIT |
| Graphile Worker | PostgreSQL | TypeScript, for Node.js | ”High performance Node.js/PostgreSQL job queue” | MIT |
| Celery | Usually a message broker; its RabbitMQ and Redis transports are “feature complete” | Python | ”Distributed Task Queue” | New BSD License |
| Sidekiq | Redis 7.0+, Valkey 7.2+ or Dragonfly 1.27+ | Ruby | ”Simple, efficient background jobs for Ruby” | LGPLv3; Sidekiq Pro and Sidekiq Enterprise are sold with “a commercial-friendly license” |
| Temporal | Its own server, the Temporal service | Go (the server) | “a durable execution platform” | MIT |
| Apache Airflow | Its own scheduler and workers, with PostgreSQL or MySQL listed in its requirements | Python | ”a platform to programmatically author, schedule, and monitor workflows” | Apache-2.0 |
The first five ride on a store you may already run, which is why they suit a small team: pg-boss and Graphile Worker run on the Postgres the app already has. Temporal and Airflow are different in kind. Each is a system of its own, built for processes with many steps, and each brings servers, upgrades and monitoring of its own.
My working rule for a small SaaS: if jobs are single-step and modest in volume, a Postgres-backed library or the plain jobs table is enough, and an orchestrator is a second system to run.
How to verify it
Background processing is verified with seven checks: note the shortest clock, start a task that runs well past it, close the tab, watch the job row finish, reopen and find the status and result, kill the worker mid-job and see the job recover or fail visibly, then start a batch of jobs at once.
Deliverable 6.6 of the Production Hardening Sprint is verified this way: “Run a long task beyond the normal request window and verify completion and user-visible status.” The checks below turn that line into steps you can run on a staging copy of your own app.
- 01 Write down the shortest clock on your host, from the tables on this page or the two timeout articles linked above. That number is the request window.
- 02 On staging, start a task built to run about two to three times that window, such as an export over a large set of seeded data. Pass: the request itself returns at once with a job id, well inside the window.
- 03 Close the tab as soon as the job id comes back.
- 04 Watch the job row. Pass: it moves from queued to running, progress changes, and it ends as succeeded, with a finished time later than the request window would have allowed.
- 05 Reopen the app as the same user. Pass: the status is visible, the result is delivered (a file link or the generated content), and the notification arrived.
- 06 Start another long job and stop the worker process halfway through. Pass: the job goes back to queued once its lease expires, or is marked failed with a reason. Fail: it stays running.
- 07 Start about twenty jobs at the same moment. Pass: the queue drains while normal pages stay responsive.
The sizes in checks 2 and 7, a task of about two to three times the window and about twenty jobs at once, are my working rule, not a standard.
Keep the evidence: the job rows with their timestamps, the timing of the original request, the notification, and the date of the run. One more thing to look at after check 6: if the restarted job ran twice, it must not have sent two emails or charged the customer twice. That is the idempotency work named in the jobs table section.
Where the sprint does this
In the Production Hardening Sprint, this page’s subject is deliverable 6.6: “Move long-running AI generation, exports, and bulk communication into background jobs.” Reliable job execution is 6.7, which is to “Make jobs idempotent and provide a dead-letter or failed-job path with a recovery procedure.” Retries are 6.5: “Add bounded retries with backoff to transient failures where repeating the operation is safe.” The production readiness report, 13.1, is verified this way: “Account for all 123 IDs; keep failures visible until resolved and explain genuine non-applicable items.” Building new product features or modules is outside the sprint. Hosting, paid tools, and API usage remain in your accounts. Every deliverable and its verify line is listed in the published scope.
Common questions about background jobs
How to run something as a background process?
On a server, end a shell command with & to run it in the background, or let a service manager keep a long-lived process running; in a web app, write a job row and let a separate worker run it. The shell route suits one-off scripts. A process started from a terminal on a server does not survive a redeploy (my reading), so in a web app the job-and-worker pattern on this page is the answer.
What does it mean the request timed out?
It means one side stopped waiting before the other answered. The work may still have finished on the server after the client gave up, which is why a blind retry can do it twice (my reading); making a repeat safe is what an idempotency key is for.
How to fix a request timed out?
Find which clock fired, using the status code and the platform’s log, then either make the work faster or stop doing it inside the request. For Vercel functions and for AI calls, the Vercel function timeout and OpenAI API timeout articles linked earlier walk through each case.
Is there a way to keep an app running in the background?
For a web app, yes: the server keeps running after the user closes the tab, so the real question is which process owns the task, and the answer should be a worker rather than the page. Keeping a phone or desktop app open in the background is a different subject, set in the device’s own settings.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase