The week a customer asks you to commit to a response time, the average stops being useful. Tail latencies are the slow end of the distribution: p95 is the time 95 percent of requests beat, p99 the time 99 percent beat. A page making 20 independent backend calls sees at least one slower than p95 on 64 percent of loads.

Tail latencies: what they are, and why the average hides them

Tail latencies are the response times of the slowest requests, read at high percentiles such as p95 and p99. The average hides them: 98 requests at 100 milliseconds and two at 10 seconds give a mean of 298 milliseconds, a figure no request experienced, while p99 shows the 10-second wait two real users sat through.

That example is what tail latency means in practice. Response times are not a bell curve. Most requests cluster near the fast end, and a minority stretch far to the right; that stretch is the tail. Every page load, API call and background job feeds the same distribution, which is why tail latency is the vocabulary every other number in web performance optimization is written in.

The mean sits where almost no request lands: nobody in the example waited 298 ms. The median does not notice the slow two at all, and neither does p95. Here are the same 100 requests read five ways (my arithmetic, easy to redo yourself):

ReadingValueWhat it tells you
Mean (average)298 msA number no request experienced
p50 (median)100 msThe typical request; blind to the slow two
p95100 msStill blind: only 2 of 100 are slow
p9910,000 msThe slow requests show up
Max10,000 msThe single worst request

For a founder, tail latency means the wait your unluckiest users sit through, the part an average buries. A widely cited paper on it is “The Tail at Scale”, Jeffrey Dean and Luiz André Barroso’s 2013 article in Communications of the ACM, which says “It is challenging to keep the tail of the latency distribution low for interactive services as the size and complexity of the system scales up or as overall utilization increases.”

Long tail latency: where the slow requests come from

Long tail latency is the stretched right side of the response-time distribution, and on a small app on managed hosting it has 7 usual causes: cold starts, memory pressure, a query missing an index, an outbound call with no deadline, a cache miss, a full connection pool and a saturated instance. The biggest accounts tend to feel it first.

“Long tail” and “long-tail latency” name the same thing as tail latency, with the shape of the curve in the name. The seven causes below are my reading of where the tail comes from on an app like that; the fix for each is a topic of its own, so the table names it rather than repeating it.

CauseWhat it looks like on a small appWhat to read up on
Cold start on a functionThe first request after a quiet spell is slow; the next ones are fastLambda concurrency and serverless limits
Memory pressure and garbage-collection pausesRequests slow down as memory fills, until the process is killed and restartswhat an OOM kill is
A query that scans because an index is missingSlow only for the accounts with the most rowshow to analyze a Postgres query
An outbound call with no deadlineThe route waits as long as the other service doeshow to set a timeout on fetch
A cache miss on a path that is usually cachedA normally fast page is slow on the first load after the cache emptiescaching strategies for web applications
Waiting for a database connection when the pool is fullRequests queue for a free connection before any query runswhy an AI app stalls at a hundred users
An instance that is simply saturatedEverything slows at peak traffic and recovers when it passeshorizontal scaling vs vertical scaling

The pool row has a signature of its own, which the stall article in that row describes. The paper’s own list of why a component’s response time varies is longer: shared resources, background daemons, global resource sharing, maintenance activities such as periodic garbage collection, and multiple layers of queueing, plus hardware trends such as power limits and energy management. Who lives in the tail on a small app is, by my reasoning, predictable: the largest accounts, because their lists are longest, their queries touch the most rows and their pages carry the most items.

Why it matters for a small app: the fan-out arithmetic

One slow request in twenty sounds rare. A page rarely makes one request, though. A dashboard that loads a user, their projects, their billing state and a dozen widgets fans out into many backend calls, and the page is at least as slow as the slowest of them.

The arithmetic is short. If each call is independently slower than p95 five percent of the time, the chance that at least one of n calls is slower is 1 minus 0.95 to the power n. For p99, use 0.99.

Backend calls per pageChance at least one is slower than p95Chance at least one is slower than p99
15%1%
523%5%
2064%18%
5092%39%

The method is Dean and Barroso’s. In the paper’s fan-out example, a server with a one-second 99th percentile slows one user request in 100 when it works alone, yet “If a user request must collect responses from 100 such servers in parallel, then 63% of user requests will take more than one second”. My reading is that the independence assumption in the table is generous: slow calls tend to cluster, because they share a database, a pool and an instance, so real pages can fare worse.

What it means in practice: on a page that fires twenty requests, the rare slow call is the normal experience, and cutting the number of calls per page is a tail fix in its own right. The article on what API response time your app actually needs shows the same effect from the other side, in its request-waterfall section.

For context from my own audits: across the third-party apps I reviewed, the Performance & Scale pillar averages 53.3 out of 100, scored on 21 of the 21 third-party apps. Those 21 are the public and held-out third-party apps I audited in June and July 2026, a selected set rather than a sample of AI-built apps in general.

How it works: the percentiles, the calculation, the four ways the number lies, and the target

The four parts below go in the order you need them: what each percentile means, how to compute one, how the number misleads, and which number to promise. The definitions and behaviors come from the sources I name in each part, and the arithmetic is the page’s own, not a test I ran.

What is p95 latency, next to p50, p90 and p99

P95 latency is the response time that 95 percent of requests beat, over a stated window and for a stated endpoint. It is the usual target for a small app because it stays stable at modest traffic. P99 shows what the unluckiest request in a hundred sees, and it needs more volume before it means anything.

Put plainly, p95 means this: sort the durations of every request to one endpoint over one window, and p95 is the value that 95 percent of them were faster than. A percentile without its window and its endpoint is not a measurement. “p95 is 400 ms” says nothing until you know it was the checkout route over the last seven days.

Here is what p50, p90, p95 and p99 latency each tell you, next to max. The “good for” and “hides” columns are my reading.

PercentilePlain meaningWhat it is good forWhat it hides
p50The typical request: half were fasterSpotting a general regressionEverything slow
p90The time 9 requests in 10 beatA first look at the tailThe slowest tenth
p95The time 19 requests in 20 beatTargets and alerts on a small app: stable with modest trafficThe slowest twentieth
p99What the unluckiest one in a hundred seesThe tail itself, once the route has volumeLittle, but it is unstable on a small sample
MaxOne requestFinding the outlier to investigateUseless as a target: one bad request moves it

You’ll also see p99.9, which the paper reports in its benchmarks of Google’s BigTable service; it is a large-system number that needs far more traffic than most small apps see on one route.

How to calculate p95 from your own logs or database

P95 is calculated in 4 steps: collect every request’s duration for one endpoint over a window, sort them, take the value at position 0.95 times the count, rounded up, and report it with the count. Postgres does it in one query with percentile_cont. With 100 samples, only one request sits above p99, so always state the sample size.

Before any query, look at what you already have. On Vercel, open the project’s Observability section in the sidebar, pick Vercel Functions and click a route; Vercel’s docs say “Latency and breakdown by path are only available for Observability Plus users”, and Observability Plus is available on Pro and Enterprise teams. Which percentile that per-route latency view charts is not stated in Vercel’s docs. Sentry’s Backend Overview dashboard shows “p50 and p75 Duration”; a per-route p95 is not stated on that docs page. On either, read the chart’s label before you quote the number, and compute p95 yourself when the label names another percentile.

To do it yourself, the method is the nearest-rank one:

  1. Collect the duration of every request to one endpoint over one window, failed and timed-out requests included.
  2. Sort the durations from fastest to slowest.
  3. Multiply the count by 0.95 and round up; that position in the sorted list is p95.
  4. Report the value with the count and the window beside it.

If your requests are logged to a Postgres table, PostgreSQL’s ordered-set aggregates do the sort and the pick in one query. percentile_cont “will interpolate between adjacent input items if needed”, while percentile_disc returns “the first value within the ordered set of aggregated argument values whose position in the ordering equals or exceeds the specified fraction”, so it always returns a duration that was actually logged. The query below assumes a table named request_log with a route text column, a duration_ms number column and a started_at timestamp column; rename them to match yours.

SELECT route,
       count(*) AS requests,
       percentile_cont(ARRAY[0.5, 0.95, 0.99])
         WITHIN GROUP (ORDER BY duration_ms) AS p50_p95_p99
FROM request_log
WHERE started_at >= now() - interval '7 days'
GROUP BY route;

Keep failed and timed-out requests in the set. Leaving them out flatters the tail, and the response-time guide linked above makes the same rule. There is a Postgres detail here too: these aggregates “ignore null values in their aggregated input”, so a timed-out request logged with an empty duration silently drops out; log its duration as the time it waited.

If your numbers come from a metrics system instead of raw rows, they are estimates. In Prometheus, “the calculation of quantiles from the buckets of a histogram happens on the Prometheus server using the histogram_quantile() function”, and for classic histograms the “Error is limited by the width of the bucket the quantile is located in”. Bucket edges set your accuracy; Prometheus on histograms and summaries covers the choice, which rests on the difference between the Prometheus metric types.

Now the sample-size floor, as arithmetic. With 100 samples, p99 is the 99th value, so one slow request more or less decides what it shows: in a set of 99 fast requests and one 10-second request, p99 reads 100 ms and the slow request vanishes from the figure, while a second slow request sends p99 to 10 seconds, as in the example at the top. With 20 samples, p95 is the 19th value, again one request from the top. My working rule of thumb is to state the count every time and to distrust a percentile resting on fewer than a few hundred requests. A low-traffic app widens the window, a week or a month, rather than trusting a small sample.

Four ways a percentile number lies

A percentile number misleads in 4 common ways: percentiles averaged across servers or hours, a load tool that stops sending while the server stalls, timing taken only inside the server, and one figure for the whole app. Each one makes the result look better than what users experience, which is why each survives review.

The mistakeWhy the number comes out too goodWhat to do instead
Averaging percentiles (three servers’ p95s, or 24 hourly p95s)The average of p95s is not the p95 of anything; one bad hour is diluted by the calm onesRecompute from the raw durations, or merge histograms that share bucket boundaries and take the quantile of the merged buckets
Coordinated omission in a load testThe tool waits for each response before sending the next request, so it stops sending while the server stalls and under-counts the stallUse a tool or mode that sends on a fixed schedule and records latency from the intended send time
Timing only inside the serverQueueing in front of the app, cold starts, TLS and the network happen before your code starts the clockCompare with a measurement taken from outside or from the browser
One number for the whole appCheap, frequent endpoints dominate a global p95 and hide the checkoutReport per journey

The first two rows lean on the tools’ own docs, and rows three and four are my reading. For averaging, Prometheus’s own example is a service running on several instances, whose summaries each precompute a 95th percentile; there, “averaging the quantiles yields statistically nonsensical values”, while with histograms “the aggregation is perfectly possible with the histogram_quantile() function”, for classic histograms “provided there are no changes in bucket boundaries”. For the load-test row, the wrk2 README on coordinated omission, written by Gil Tene, explains that “Since each connection will only begin to send a request after receiving a response, high latency responses result in the load generator coordinating with the server to avoid measurement during high latency periods.” Its fix: wrk2 “measures response latency from the time the transmission should have occurred according to the constant throughput configured for the run.” The method for running the test itself belongs to load testing a web application.

Rows one and four can land together, as in this case. A founder reports one response-time figure for the whole app: the average of a dashboard’s hourly p95s. The largest customer then says checkout is slow. The checkout route’s own p95, computed from raw durations over the same window, sits far above the app-wide figure, because the cheap, frequent routes carried the global number and an average of hourly percentiles is not a percentile of anything. The lesson I take from a case like this: a percentile only tells the truth with its window, its count and the way it was measured written beside it.

What number to hold the app to: the percentile, the window and the count

A latency target, by my working rule, names its percentile, not only its threshold: p95 for a route with modest traffic, p99 once the route has the volume to hold it steady, each over a stated window with the request count beside it. The threshold itself is a product choice for each action, set from what the route measures today.

The response-time guide linked above covers the threshold, the human-perception limits and the shape of the target sentence. What it leaves to you is which percentile that sentence names and how the window and count are chosen. My working rule: pick the three journeys that carry money or trust (sign-in, the main task, checkout), measure each for about two weeks, name p95 unless the route has the volume for p99, write the window and the request count into the target, and set the threshold a little above today’s measured value so a regression trips it.

Browser-side Core Web Vitals are also judged at a percentile of page loads, with thresholds of their own, which the Core Web Vitals guide covers. A target you promise a customer in writing becomes a service level, so read up on what SLAs are before you sign one. The same target sentence belongs in your capacity statement, which is part of what a capacity plan is, and it drives the alert, which makes how to alert on error rate spikes and other thresholds the next thing to settle.

The paper also describes hedged requests, which “issue the same request to multiple replicas and use the results from whichever replica responds first”, and “tied requests”, its name for requests “where servers perform cross-server status updates”. I don’t advise either on a small app: duplicating a request is unsafe for writes and adds load, and the paper itself notes that naive implementations “typically add unacceptable additional load”.

How to check your own app

Tail latency is checked 6 ways: percentiles exist per route with counts, none is an average of percentiles, one journey is timed from outside and compared, the load test reports p95 and p99 with its pacing mode, a p95 alert has fired once on purpose, and after a fix the same query shows the tail moved.

Each check below is written so a correct setup passes and a broken one fails, with the evidence to keep.

  1. 01 Per-route percentiles exist. p50, p95 and p99 for the three journeys, over at least seven days, with the request count beside each. Pass: every figure has its count and its window. Keep: the query and its output.
  2. 02 No figure is an average of percentiles. For each percentile on a dashboard or report, find how it is computed, from the tool docs or the query. Pass: each comes from raw durations or merged histograms. Keep: a note of the source for each figure.
  3. 03 One journey is timed from outside. Time it against the deployed production site, never the local dev server, with a synthetic check or the network panel of a desktop browser over several runs, and compare it with the server figure for the same route and window. Fail: a gap larger than network and TLS explain. Keep: both numbers.
  4. 04 The load test reports p95 and p99 per journey. The report states the pacing mode (a fixed request rate, or each connection waiting for a response before it sends again), the workload, duration, environment and error rate, before and after changes. Pass: all of it is written down. Keep: the report.
  5. 05 The p95 alert for the main journey has fired once on purpose. In a copy of the alert, set the threshold below the current p95, wait out the evaluation window your provider documents, confirm it fired and reached the person, then restore it. A threshold just above the current value is not a test. Keep: the alert record.
  6. 06 After a fix, the tail moved. Run the same query over the same window length and compare p95 and p99, with counts. Fail: the average improved while p99 did not. Keep: the before and after output with dates.

In the Production Hardening Sprint, deliverable 9.2, Load test and bottleneck fixes, is verified this way, close to check four: report the workload, duration, environment, concurrency, latency, and error rate before and after changes.

Where the sprint fits

No deliverable in the sprint is named tail latency. The four closest, by their titles in the published scope, are Load test and bottleneck fixes (9.2), which is to simulate concurrent users, identify the first bottlenecks, fix them, and rerun the workload; Written capacity statement (9.6), which is to document measured concurrent capacity on the current infrastructure and the changes needed to plan for five times that workload; Operational threshold alerts (8.3), which alert on error spikes, latency, connection pressure, and queue backlog; and Unified operations dashboard (8.8), which creates one view of uptime, error rate, latency, signups, and revenue for the application’s relevant services. Hosting, paid tools, and API usage remain in the client’s accounts.

Common questions about latency percentiles

Which is better, p95 or P99?

Neither on its own. P95 is the better choice for targets and alerts at modest traffic because it holds steady, and p99 earns its place once a route has the volume to keep it stable and the slow one percent includes your most valuable accounts. Report both, each with its request count, and let the traffic decide which one the target names.

Is P50 or P90 better?

They answer different questions, so neither replaces the other. P50 shows the typical request and catches a general regression; p90 is where the tail starts to show. A report that gives only p50 is hiding the slow requests, which is where complaints come from.

Is 40 ms of latency bad?

Not for judging an app: as a network ping, 40 ms describes a home or office connection, not how long your app takes to answer a request. For an app, the number that matters is the p95 of each journey timed from the user’s side, which includes the network and everything the server does. What value that p95 should stay under is a product choice per action, and the response-time guide covers how to set it.

How do I fix network latency?

By my working rule, host the app and the database in the same region, close to your users, serve static assets from a CDN, and keep connections to the database and other services warm rather than opening a new one per request. Before any of that, check that the network is the problem: if the server’s own p95 for the route is the large number, the network is not.