A serverless app does not slow down when it runs out of room: it starts refusing requests. Lambda concurrency is the number of requests your functions are handling at the same instant, and once an account reaches its quota, 1,000 per Region by default, a synchronous request gets a 429 throttling error instead of a place in a queue.
What is Lambda concurrency? How to work out yours from two numbers
Lambda concurrency is the number of requests your functions are processing simultaneously, because each execution environment serves one request at a time. To estimate it, multiply average requests per second by average duration in seconds. A slow function can use more concurrency than a busy one, so duration is the number to fix first.
It is one ceiling among several in web performance optimization, and the one that turns a slow app into an app that says no. AWS’s guide to Lambda scaling and concurrency puts the model plainly: concurrency is the number of in-flight requests a function is handling at the same time, and for each concurrent request Lambda provisions a separate instance of the execution environment. Ten requests in flight means ten environments. In AWS Lambda, concurrency is counted in environments busy at one instant, not in users and not in requests per second. The formula AWS gives is: concurrency = (average requests per second) × (average request duration in seconds).
Here is that formula worked for a small SaaS with three endpoints. The traffic figures are an illustration I picked to show the shape, not a measurement of any app.
| Endpoint | Requests per second at peak | Average duration | Concurrency |
|---|---|---|---|
| Sign-in (a fast database read) | 20 | 0.1 seconds | 2 |
| Checkout (calls a payment provider) | 2 | 1.5 seconds | 3 |
| AI generation (waits on a model) | 1 | 20 seconds | 20 |
The AI endpoint gets the least traffic and holds 20 of the 25 environments, because each call stays open for 20 seconds. Two things follow from the formula, in my reading. Halve a function’s duration and you halve its concurrency. And a dependency that slows down raises your concurrency without a single new user arriving.
Concurrency differs from requests per second, and AWS caps both. Across all functions in an account, Lambda enforces a requests-per-second limit equal to 10 times the account concurrency, so the default 1,000 allows up to 10,000 requests per second. How many concurrent requests Lambda can handle for you is therefore two numbers: 1,000 in flight across a Region’s functions by default, and 10 times that per second.
Concurrent executions are the same count under the name CloudWatch uses. The ConcurrentExecutions metric is the number of function instances processing events, and AWS says to read it with the Max statistic to see how close you are to the limit. That graph is where AWS Lambda concurrent executions show up after the fact. Filtered to one function, it shows AWS Lambda function concurrency; for the whole Region, it is the total you compare against the quota.
Why it matters when your functions run on Vercel or Netlify: the same limits under other names
If your back end runs as Vercel or Netlify functions, you may never open the AWS console. The model is the same in my reading: an instance per request in flight, a platform ceiling on instances, and a hard stop on duration. The table puts four platforms side by side, each cell from that vendor’s own docs.
| Platform | What the concurrency ceiling is called | The time ceiling | The error a user sees | Where the vendor documents it |
|---|---|---|---|---|
| AWS Lambda | Concurrent executions, 1,000 per Region by default | 900 seconds (15 minutes) for a standard function | A throttling error (429 status code) on a synchronous request | The Lambda quotas and scaling pages |
| Vercel Functions | ”the concurrent execution limit” | Set per plan (see the Vercel link below) | 503 FUNCTION_THROTTLED; a timeout returns 504 FUNCTION_INVOCATION_TIMEOUT | Vercel’s error reference |
| Netlify Functions | Not stated in Netlify’s functions docs | 60 seconds synchronous, 30 seconds scheduled, 15 minutes background | Throttling error not stated in Netlify’s functions docs; a background function answers at once with a 202 | Netlify’s functions configuration and background functions pages |
| Supabase Edge Functions | Not stated on Supabase’s limits page | A worker’s wall clock: 150s on Free, 400s on paid plans; 2s of CPU time per request | 504 Gateway Timeout after a 150s request idle timeout | Supabase’s Edge Functions limits page |
Limits checked 30 September 2026.
The AWS row comes from the Lambda quotas page. The Vercel cells come from Vercel’s FUNCTION_THROTTLED reference, which says the error “occurs when your Vercel Functions exceed the concurrent execution limit, often due to a sudden request spike or backend API issues”; Vercel’s errors index gives it a 503 status. The timeout cell is the other Vercel error: FUNCTION_INVOCATION_TIMEOUT, a 504 Gateway Timeout when a function exceeds its duration budget. Netlify’s background functions “don’t need to complete before a visitor can take next steps on your site”, and they are available on Credit-based plans, including Free, Personal, and Pro, and on Enterprise plans. For Supabase, Supabase’s Edge Functions limits also cap memory at 256MB.
Three neighboring questions each have their own article. Vercel’s plan-by-plan durations and maxDuration are in Vercel function timeout limits and fixes. Which clock hangs up first on an AI call, the SDK’s default or the host’s limit, is a comparison of its own: the OpenAI API timeout and the host limit that hangs up first. And which ceiling each kind of platform tends to hit first under a crowd is in why an AI app slows down under concurrent users and stalls at 100.
In my June and July 2026 audits, 13 of the 21 third-party apps had no rate limiting on their most expensive endpoint. Those 21 are apps I chose to audit, a selected set rather than a random sample, so the figure says what I met, not how common this is across AI-built apps. My reading of what that means on a serverless host: a stranger’s burst becomes your concurrency, and with nothing in front of the function, the platform quota is the only brake. Putting a limit in front of the expensive endpoint is rate limiting in an API.
How it works: the limits, the time ceiling, the database behind it, and the fix order
The first two parts below are Lambda’s own limits; the third sits behind the functions, in the database. On a small app they tend to show up in a different order, in my reading: the database first, then the timeout in front of the function, and the account quota last.
Lambda concurrency limits: the account quota, the scaling rate, reserved and provisioned
Lambda concurrency limits start with an account quota shared by every function in a Region, 1,000 by default. At the quota, a synchronous request fails with a 429 throttling error. Reserved concurrency sets aside and caps one function’s share at no charge. Provisioned concurrency keeps environments initialized and bills while it is configured.
Four controls decide how much of that quota a function gets. The table gives each one in AWS’s terms.
| Control | What it does | The default | Who pays, and when to use it |
|---|---|---|---|
| Account concurrency quota | One pool shared by every function in the Region | 1,000; new accounts start lower and AWS raises them automatically based on usage | Raising it might add cost, AWS warns; request it last, with your numbers |
| Concurrency scaling rate | How fast each function can add execution environments | 1,000 environments every 10 seconds, per function | Not a setting; AWS says you usually don’t need to worry about it |
| Reserved concurrency | Sets both the maximum and minimum concurrent instances for one function; no other function can use them | Not set; you can reserve up to the unreserved account value minus 100 | No additional charge; for a function that can starve the others or overload a database |
| Provisioned concurrency | Keeps pre-initialized execution environments ready to reduce cold start latencies | Not set | Billed for the concurrency configured and for how long, plus requests and duration; for measured cold starts on a user-facing function |
The account quota row is the AWS Lambda concurrency limit most people mean, and it can be raised through Service Quotas; the quotas page says it can go up to “Tens of thousands”. Both reserved and provisioned concurrency count toward it, so every unit you allocate comes out of the pool the other functions share. Rows 3 and 4 are the difference between reserved and provisioned concurrency in AWS Lambda: one is a slice of the quota with a ceiling, the other is warm environments you pay for. Configuration steps are in AWS’s reserved concurrency guide, and the billing terms in Lambda pricing.
The per-function scaling rate dates from December 6, 2023; before that, scaling was at the account level, “up to 3000 concurrent executions in the first minute, followed by 500 concurrent executions every minute afterwards”. A burst of 3,000 in the first minute and 500 a minute after it describes that earlier model.
What happens when the Lambda concurrency limit is reached depends on how the function was called. A synchronous caller, such as API Gateway waiting on a response, gets the failure straight away: additional requests “fail with a throttling error (429 status code)”. For an asynchronous invocation, Lambda returns the event to its queue and attempts to run the function again for up to 6 hours by default; when an event expires or fails all processing attempts, Lambda discards it, and an on-failure destination is how AWS says to capture records of failed invocations. For an SQS queue, Lambda backs off and keeps retrying the message until its timestamp exceeds the queue’s visibility timeout, “at which point Lambda drops the message”.
Throttles are easy to miss. AWS’s Lambda metrics reference says throttled requests “don’t count as either Invocations or Errors”; they appear in the Throttles metric. So an app can sit at the AWS Lambda concurrent execution limit with a flat Errors graph while users get refused, which is why the checks below put the alarm on Throttles.
AWS Lambda max runtime: 15 minutes for the function, far less for an HTTP request
AWS Lambda’s max runtime for a standard function is 900 seconds, 15 minutes, and a new function’s default timeout is 3 seconds. A web request behind API Gateway gets far less: an HTTP API’s integration timeout is 30 seconds and cannot be raised. Work that needs minutes should return an id at once and finish in the background.
The function timeout is set in 1-second increments up to 900 seconds. That 900 seconds is the AWS Lambda max execution time for a standard function, and the quotas page lists function configuration quotas like it as ones that “can’t be changed” except as noted. The one exception it notes is Lambda Managed Instances, where asynchronous invocations and event source mapping invocations (except Amazon MQ and Amazon DocumentDB) support up to 5,400 seconds, 90 minutes. When a function reaches its timeout, “Lambda stops the function invocation”. AWS’s troubleshooting page shows a timeout during the Init phase as Task timed out after 3.00 seconds.
The AWS Lambda execution time limit a user feels is usually set in front of the function. Four clocks run on one web request:
| Layer | Its timeout | Default versus maximum | What the caller sees when it is hit |
|---|---|---|---|
| Lambda function | The function’s Timeout setting | 3 seconds by default, up to 900 seconds for a standard function | Lambda stops the invocation |
| API Gateway | The integration timeout | HTTP API: 30 seconds, cannot be increased. REST API: 50 milliseconds to 29 seconds; Regional and private APIs can go higher, which might require a lower Region-level throttle quota | Not stated on the API Gateway quotas pages |
| An outbound call inside the function | Whatever timeout your HTTP or SDK client sets | Set by the library, often long (my reading) | Nothing yet: the function keeps waiting and holds its environment (my reading) |
| The client | How long the browser or mobile app waits | Set by your front end (my reading) | A spinner, then often a retry that adds a request (my reading) |
Those front-door figures come from the API Gateway quotas page for REST APIs and its HTTP API counterpart. A function URL’s own timeout is not stated in AWS’s function URL docs, checked 30 September 2026. So the maximum execution time for a Lambda function that a user waits on behind an HTTP API is 30 seconds, whatever the function timeout says; the Lambda runtime limit that matters for a page load is the front door’s. The function timeout itself is set in AWS’s function timeout guide.
My working rule follows from the table: no user-facing request should depend on a function running for minutes. Long work returns an id at once and finishes as a job, which is how to run long tasks in the background. And every outbound call inside a function needs its own shorter deadline, so one slow provider cannot hold an environment until the ceiling; that is also a concurrency fix, and the pattern is how to set a timeout on fetch. Because a standard function’s 900 seconds cannot be raised, work that outgrows it has three ways out in my reading: split it into steps, put it on a queue with a worker, or move it to another compute service. The Lambda execution limit on a single invocation is not something to design up to.
The limit that often bites first: database connections behind the functions
The serverless limit that often bites first is the database, not Lambda. Each execution environment that talks to the database holds a connection of its own, so function concurrency becomes database connections, and a small managed Postgres often accepts fewer than Lambda’s default quota allows. A pooler, a reused client and a reserved concurrency cap are the three fixes.
The arithmetic is short. Supabase’s compute docs list 60 database max connections and 200 connection pooler max clients for the Micro and free Nano sizes, and 90 and 400 for Small; the connection figures are recommended values that can be customized. Against a Lambda quota of 1,000, a Micro database runs out of direct connections long before Lambda runs out of environments. It can be worse than one connection per environment: Supabase notes that Postgres.js defaults to 10 connections, “10 connections for every warm instance of your function”.
The three fixes, in the order I’d apply them (my working rule):
First, put a pooler between the functions and the database. On Supabase that is the transaction-mode pooler, which Supabase recommends for serverless functions; transaction mode does not support prepared statements, so turn them off in your connection library. On AWS the equivalent is Amazon RDS Proxy, which pools and shares connections and queues or throttles the ones it cannot serve at once; it must be in the same VPC as the database, and the proxy “can’t be publicly accessible”.
Second, create the client once, outside the handler, so warm invocations reuse it. AWS calls it best practice to use the INIT phase “to set up database connections”, and Supabase’s settings for a serverless function are a pool of 1 with prepared statements off:
import postgres from 'postgres'
// created once per execution environment, reused by warm invocations
const sql = postgres(process.env.DATABASE_URL, { max: 1, prepare: false, ssl: 'require' })
export const handler = async () => {
const rows = await sql`select id, name from projects limit 20`
return { statusCode: 200, body: JSON.stringify(rows) }
}
Third, cap reserved concurrency on every function that connects, below what the pooler or database can take. AWS describes this use directly: reserved concurrency “can be used for limiting concurrency to prevent overwhelming downstream resources, like database connections”. The throttle then lands on the function, where the queue or the caller can retry, instead of on the database, where every user fails at once.
Pooling in depth is connection pooling in Postgres. The error strings and the outage steps are in the connection pool exhausted errors a serverless app sees, and the same problem seen as slow responses is in what API response time your app actually needs. Memory is the other ceiling inside the function: it is set between 128 MB and 10,240 MB, and AWS notes that some database connection and logging libraries might keep data between warm invocations, so their memory grows and the function might run out of memory. How a process gets ended for that is what an OOM kill is.
When you are throttled: what to change, in order
A throttled Lambda app is fixed in six steps, in this order: confirm the throttle in the metrics, find the slowest function, shorten it, put bursty work behind a queue with a capped consumer, reserve concurrency for functions that can starve others, and only then request a higher quota. Provisioned concurrency fixes cold starts, not throttling.
- 01 Confirm it is a throttle, not an error. In CloudWatch, read Throttles and ConcurrentExecutions (Max statistic) for the function and for the Region. On Vercel, search the function logs for FUNCTION_THROTTLED from the platform table above (Netlify's functions docs name no throttling error). Confirmed when the Throttles graph rises at the same minutes users saw failures.
- 02 Find the slow function, not the busy one. Sort functions by average duration, then multiply each by its requests per second. Confirmed when one function accounts for most of the concurrency at the peak.
- 03 Make that function shorter. Give every outbound call its own deadline, move work that the user does not wait for out of the request, and cache what repeats. Confirmed when its average duration drops and its concurrency drops with it.
- 04 Move long or bursty work behind a queue and cap the consumer with the SQS maximum concurrency setting, so a spike becomes a backlog instead of a throttle. Confirmed when a burst shows up as queue depth, with Throttles flat.
- 05 Set reserved concurrency on the function that can starve the others, and on every function that connects to the database. Confirmed when the capped function throttles alone and sign-in and checkout keep answering.
- 06 Only then ask for a higher account quota through Service Quotas, with the calculation from the first section and a load test in the request. Confirmed when Service Quotas shows the new value for the Region.
The order is my working rule: duration first, the quota request last. Maximum concurrency on an SQS source takes a number between 2 and 1,000, costs nothing to configure, and should not be set higher than the function’s reserved concurrency; the steps are in the SQS maximum concurrency setting. For step 6, AWS’s guide to requesting a concurrency increase asks for your expected average and peak requests per second, runtime duration, memory size, invocation type and event source, plus “Load test results”, and suggests asking at least two weeks before you need it. Provisioned concurrency is left off the list on purpose: it reduces cold start latencies and bills for as long as it is configured, and neither of those helps a function that is out of quota.
Step 5 from the command line, with AWS’s own example value:
# reserve 100 concurrency units for one function (AWS's example)
aws lambda put-function-concurrency --function-name my-function \
--reserved-concurrent-executions 100
A case shows why step 2 comes before any quota request. Picture your sign-in, checkout and AI generation endpoints running as Lambda functions in one Region, with no reserved concurrency set. Then the AI provider slows down. Traffic stays the same, but the generation function’s duration grows, so its concurrency grows with it, because concurrency is requests per second times duration. The account quota is shared by every function in the Region, so sign-in and checkout start answering with throttles although their own traffic did not change. Reserved concurrency on the generation function would have capped its share, and a deadline on its outbound call would have bounded its duration; a quota increase would only have moved the ceiling. The function to fix is the one whose duration grew, not the one whose traffic did.
Sometimes the honest answer is that the workload no longer fits functions: steady heavy traffic, long jobs, big memory. Choosing between a bigger machine and more of them is horizontal scaling vs vertical scaling.
How to check your own app
Serverless limits are checked six ways: the account quota is written down, peak concurrent executions sit well under it, a throttle alarm has fired once, each function’s timeout sits next to its p99, the connection limit exceeds the reserved concurrency of the functions that connect, and a staged load test at peak shows zero throttles.
Start in the console, no terminal needed. Open the Service Quotas console, choose AWS Lambda, choose View quotas, and select Concurrent executions for the Region the app runs in. If the number is below 1,000, the account likely still has the reduced quota AWS gives new accounts. Then run the six checks. Each one passes on a correct setup and fails on a broken one.
- 01 The account quota and the unreserved remainder are written down with the date. Pass: both numbers and the date in your notes. Fail: nobody knows the quota. Evidence: a screenshot of the Service Quotas page.
- 02 The peak of the Region's ConcurrentExecutions over the last two weeks sits under about half the quota, or there is a written plan for the gap. Pass: the Max graph and the plan if needed. Fail: peaks near the quota with no plan. Evidence: the graph with its date range.
- 03 An alarm on Throttles above zero exists and has fired once on purpose. On a staging copy of the function, set reserved concurrency to 0, invoke it, then remove the limit. Pass: the alarm message with its time. Fail: no alarm, or an alarm nobody received. Evidence: the message.
- 04 Every function's configured timeout is listed next to its p99 duration, and no user-facing function's timeout is longer than its front door's. Pass: a table with both columns. Fail: a user-facing function set to minutes behind a 30-second gateway. Evidence: the table.
- 05 The connection limit of whatever the functions connect to (the pooler's client limit if there is one, otherwise the database's) is written next to the sum of reserved concurrency of the functions that connect, and the sum is smaller. Fail: a connecting function with no reserved concurrency. Evidence: both numbers.
- 06 A staged load test at the expected peak shows zero throttles and zero refused connections, and records the concurrency it reached. It runs in a separate account, or against a staging function capped by its own reserved concurrency, never uncapped on the quota production shares. Pass: the report. Fail: throttles, refused connections, or a test run in the production account uncapped. Evidence: the report and its date.
Check 2’s “about half” is my working rule, not an AWS figure. The alarm test in check 3 works because setting reserved concurrency to 0 “stops your function from processing any events until you remove the limit”. The load test stays off production’s share because the account quota is one pool for every function in a Region, so an uncapped test in the same account can throttle live users. Keep the evidence together: the quota screenshot, the metric graphs, the alarm message, the load test report, each with its date.
Setting up the alarm itself is part of AWS logging and monitoring. Reading p99 next to the timeout is tail latencies and p99. Designing the test so it measures something is load testing a web application, and the checks for the rest of the stack are in scaling web applications.
Where the sprint fits
The Production Hardening Sprint covers this ground in five deliverables. In 9.2, load test and bottleneck fixes, we simulate concurrent users, identify the first bottlenecks, fix them, and rerun the workload, and we report the workload, duration, environment, concurrency, latency, and error rate before and after changes. In 9.6, written capacity statement, we document measured concurrent capacity on the current infrastructure and the changes needed to plan for five times that workload, and we link capacity claims to load-test evidence and identify untested projections as projections. Deliverable 6.4 sets deliberate timeouts for AI, email, payment, and other external calls. In 6.6 we move long-running AI generation, exports, and bulk communication into background jobs. Deliverable 8.3 alerts on error spikes, latency, connection pressure, and queue backlog. Your app’s current framework and hosting setup are our starting point; we refactor or replace components where the production work requires it. Hosting, paid tools, and API usage remain in your accounts. Each deliverable and how it is verified is listed in the published scope.
Common questions about Lambda limits
How much does Lambda provisioned concurrency cost?
Lambda provisioned concurrency is billed for the amount of concurrency you configure and for how long it is configured, rounded up to the nearest five minutes, with requests and duration charged on top. AWS’s pricing examples use $0.0000041667 per GB-second on x86 in US East (N. Virginia), so one 1 GB x86 environment kept warm for a 30-day month works out to about $10.80 before request and duration charges.
The price depends on the memory you allocate and the concurrency you configure, and the Lambda free tier does not apply to functions with provisioned concurrency enabled. Prices checked 30 September 2026.
Why is the Lambda timeout 15 minutes?
The 15-minute ceiling for a standard function dates from October 10, 2018, when AWS raised it from 5 minutes “to allow for long-running functions”, as its Lambda document history records. AWS’s quotas page describes Lambda as “designed for short-lived compute tasks that do not retain or rely upon state between invocations”.
The same page now lists a longer ceiling in one case only: up to 90 minutes for Lambda Managed Instances functions invoked asynchronously or through an event source mapping, except Amazon MQ and Amazon DocumentDB.
When to use provisioned concurrency?
Use provisioned concurrency when cold starts on a user-facing function have been measured and they matter to the people using it. It keeps pre-initialized execution environments ready, which AWS describes as useful for reducing cold start latencies.
My rule is never to use it as a throttling fix: it bills while configured and counts toward the same account quota as everything else.
Why is my Lambda function getting timed out?
A Lambda function times out when its work takes longer than the timeout setting says it can run, and four causes are the ones to check first: its timeout is still the 3-second default, an outbound call has no deadline of its own, it is waiting on a database that is refusing connections, or the work is too long for a request.
AWS lists a slow response from another service among the common causes. The error to search your logs for begins Task timed out after, as in AWS’s troubleshooting example.
How do I reserve concurrency for my Lambda function?
Reserve concurrency in the Lambda console: open the function, choose Configuration, then Concurrency, choose Edit, choose Reserve concurrency, enter the amount, and save. From the AWS CLI, aws lambda put-function-concurrency with --reserved-concurrent-executions does the same.
You can reserve up to the unreserved account concurrency minus 100, because AWS keeps 100 units for functions without a reservation, and setting it to 0 throttles the function on purpose.
What are the limitations of AWS Lambda?
The limits that matter most to a small app are the concurrency quota (1,000 per Region by default), the 900-second timeout for a standard function, memory between 128 MB and 10,240 MB, invocation payloads of 6 MB each way for synchronous calls and 1 MB for asynchronous ones, a 250 MB unzipped deployment package, and /tmp storage between 512 MB and 10,240 MB.
Warm environments do keep global objects between invocations, which is what lets a database client be reused, but those are completely reset on a cold start, so anything that must survive a request belongs in a database, a queue or storage.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase