My working rule: within a month of the first paying customer, your app should tell you it is degrading before a customer does. That takes 5 threshold alerts. How to alert on error rate spikes comes first: I watch the share of requests failing over about 5 minutes, then add latency, database connections, queue backlog and spend.
What a threshold alert is
A threshold alert is a rule with 3 parts: a metric, a limit and a window. It notifies a person when the metric stays past the limit for the whole window. An uptime check says the app is down; a threshold alert says it is getting worse.
Put precisely, a threshold alert watches one number, compares it with a limit over a window, and notifies a person when the limit stays crossed for long enough. An uptime check answers a narrower question from outside the app: up or down. An error tracker’s new-issue notice fires on one event, the first time a new kind of error appears. Neither notices errors creeping from rare to common after a deploy, or a page that gets a little slower each week. Threshold alerts exist for that gradual degradation.
Every rule needs the metric, the limit and the window, plus a fourth part that matters as much: who receives it. Routing, on-call and tuning a threshold that turns out noisy belong to Slack alerting, so this page sticks to the numbers. Threshold alerts are one part of logging and monitoring for a small SaaS. They are deliverable 8.3, Operational threshold alerts, in the Production Hardening Sprint’s published scope.
On a small stack the rules live in four places it already has: the error tracker, the host’s or database provider’s metrics page, the cloud provider’s alarms (CloudWatch on AWS), and the API vendor’s usage or billing page. Nothing below needs a new monitoring platform.
What goes wrong without it
Each row is a way a live app degrades without going down. The middle column is my reading of how each one usually comes to light when nothing alerts on it.
| What degrades | How it usually surfaces with no alert | The alert on this page that catches it |
|---|---|---|
| The share of failing requests rises after a deploy | A customer writes in, and that message is the first report | Error rate over a window |
| One page gets slower | Someone complains that it is slow, or stops using it without saying why | Latency on the slow tail |
| The connection pool fills at busy times | Requests start failing with connection errors during peaks | Connection pressure |
| A worker stops and jobs pile up | Emails and exports never arrive while the site looks fine | Queue backlog: depth and oldest age |
| A retry loop or an abused endpoint runs up the API bill | The invoice | API spend |
17 of the 21 third-party apps had no error tracking or alerting: when a user hits an error, nothing records it. In the same set, 12 of the 14 AI apps had a confirmed denial-of-wallet path, where a stranger or free account can burn the owner’s paid AI or compute bill without limit. Both counts come from the third-party apps I audited in June and July 2026: 11 public vibe-coded apps audited exhaustively across all 12 pillars, and 10 held-out apps audited blind, with the 14 being the AI apps in that group. That is a selected set of audited apps, not a random sample or a rate for all AI-built apps.
A paid API with no limit in front of it is the reason spend sits on this list beside four engineering signals.
A queue backlog growing unnoticed
A queue backlog grows unnoticed because the app keeps answering while delayed work piles up. Two numbers show it: how many messages are waiting and how long the oldest one has waited. Alarm on both, because the age tells you how late the work already is for the customer waiting on it.
Take a founder’s app that hands slow work, such as welcome emails and exports, to an SQS queue, with a separate worker process taking jobs off it. When the worker stops, the site keeps answering and every request succeeds, so an uptime check stays green, while messages wait and the age of the oldest unprocessed message keeps climbing. That climbing age is what an age alarm watches, and the lesson is to watch how long the oldest job has waited, not only whether the site answers.
Running the workers, retrying failed jobs and handling dead letters is its own job: how to run long tasks in the background. A scheduled job that stops silently needs a different check again, the heartbeat check in the uptime monitoring guide.
How to alert on error rate spikes: a share of requests over a window, not a count
An error rate alert watches the share of requests that fail over a rolling window, not a raw count. My starting rule for a small SaaS is a warning above about 2 percent for 5 minutes, with a minimum number of failures so that one error at night does not wake anyone.
The number is failed requests divided by all requests over a rolling window. Failed means HTTP 5xx responses, which say the server erred, plus any failure your app reports under a 200, the code for success, such as a JSON body saying the payment did not go through. A raw count misleads in both directions: the same count is normal at peak traffic and alarming at night.
My working rule: above about 2 percent for 5 minutes is a warning that goes to the team channel, and above about 5 percent for 5 minutes wakes someone. Treat both as starting values to tune against your own traffic, not as a standard.
The low-traffic floor matters as much as the percentage. With 40 requests in the window, a single failure is already 2.5 percent, over the warning line. So I add a minimum count, for example at least about 10 failures in the window, or lengthen the window overnight. Google’s SRE workbook makes the same point with its own numbers: “if a system receives 10 requests per hour, then a single failed request results in an hourly error rate of 10%.”
Alert per critical route as well as globally: sign-in, checkout and the webhook endpoint each get their own rule. A route that fails completely but carries little traffic barely moves the global share, so a checkout that errors on every attempt can hide under a healthy total.
A “new error type after a deploy” rule is a different kind of alert. It lives on the error tracker’s issue side and fires on the first event of a new error, so keep it, but it does not replace the rate rule.
The mature form, for the app that outgrows fixed percentages, is alerting on an error budget against a service level objective (SLO). Google’s SRE workbook on alerting on SLOs says “the error budget gives the number of allowed bad events”, and that burn rate “is how fast, relative to the SLO, the service consumes the error budget.” The setup it calls the most appropriate in most cases, multiwindow, multi-burn-rate alerting, checks a long window and a shorter one, so a page fires only while the budget is still being consumed. That needs a written SLO first; the fixed rule above is the step before it.
An alert that asks for no decision teaches people to mute the channel. The logging side of that rule is to alert on decisions, not every log line, which comes with its own list of starting conditions.
How to set up Sentry alerts for an error spike
Sentry alerts on an error spike in two parts. A metric monitor tracks a threshold on one project’s errors, for one environment where Sentry offers that choice, and opens an issue when it is crossed. An alert then acts on that issue, sending it to email, or on the Team plan and above to Slack or an on-call tool.
Sentry here means the error tracker, not Tesla’s Sentry Mode. Sentry’s monitors documentation describes Monitors as the place to decide when errors and performance problems become issues, and metric monitors as a way to “track thresholds on errors, spans, logs, releases, and application metrics” using “an absolute number threshold, a percentage change, or dynamic anomaly detection.” Sentry’s alerts documentation says Alerts “take action when issues in your organization match pre-defined rules”, with actions such as email, Slack and PagerDuty. The steps below follow those pages and work from the dashboard alone.
- 01 Open Monitors, click Create Monitor and choose the metric type.
- 02 Name the monitor, select the project and, where Sentry offers it, the production environment.
- 03 Choose errors as the dataset, pick the metric, and set the interval for how often Sentry checks.
- 04 Set the threshold as an absolute number or a percentage change, then set the priority and auto-resolve.
- 05 Go to Monitors > Alerts, click Create Alert, and choose specific monitors as the source, with this monitor selected.
- 06 Choose the trigger under When, add the action under Then, and click Send Test Notification on that If/Then block to check the wiring.
Which person or channel the action points at is a routing decision that belongs with the rest of your on-call setup. Send Test Notification proves the integration is wired up, but Sentry adds that “event based triggers may behave differently”, so the drill further down still fires the real threshold.
Plan conditions shape these steps. Sentry’s pricing page lists Metric Monitors on every plan, 20 on Developer and 20 on Team, but “Alerts and notifications via integrated tools” start at Team, so on the free Developer plan the action is email. “Anomaly Detection” starts at Business, which is why step 4 uses an absolute or percentage-change threshold.
One limit, in my reading: the errors dataset counts error events, so this monitor catches a jump in the count. The share of failing requests also needs the request count, which on Sentry comes from spans, and spans arrive only once tracing is set up in the SDK. Installing the SDK and uploading source maps is its own setup job.
The other four thresholds: latency, connections, queue depth, spend
Each section below gives the number to watch, a starting threshold, the window and where a small stack sets it; how to fire each one on purpose is in the drill under the checklist. Every threshold is my working rule, a starting value to tune rather than a standard.
Latency: alert on the slow tail, not the average
A latency alert should watch the 95th percentile on key routes, because an average hides the slow tail. My starting rule is p95 above about twice its normal value, or above about 2 seconds, for 10 minutes. Set it where response times are already measured.
Google’s SRE book gives the case of a web service with an average latency of 100 ms at 1,000 requests per second, where “1% of requests might easily take 5 seconds.” The p95 figure reports the slow end directly, and a limit of twice the normal value adapts to a route that is always a little slow.
Three places can hold the rule. A Sentry metric monitor on spans works once tracing is set up in the SDK. If your host charts response time, its metrics page is the second. The third is an outside uptime monitor, and response-time thresholds on an outside check belong with the rest of that monitor’s setup. What p95 should be under heavy traffic is a question for a load test, not for the alert. The temporary slow route in the drill fires it.
Connection pressure: alert before the pool is full
A connection pressure alert fires before the database pool is full. My starting rule is active connections above about 80 percent of the maximum for 5 minutes, read from the provider’s metrics page or from pg_stat_activity. Fixing exhaustion is a separate job from alerting on it.
The number is connections in use as a share of the maximum the server accepts. On Supabase the maximum depends on compute size: Supabase’s compute and connection limits list 60 Database Max Connections on the free Nano instance, and call these “recommended values” that “can be customized via max_connections”. Usage is charted in Supabase’s Reports, where the Database connections chart (pooler connections to the database) is listed for the Free and Pro plans; the Reports docs describe the chart, not an alert on it.
On any PostgreSQL, PostgreSQL’s pg_stat_activity documentation describes a view with one row per server process, and the max_connections setting sets “the maximum number of concurrent connections to the database server.” This query returns both numbers:
select count(*) as in_use,
current_setting('max_connections')::int as max_allowed
from pg_stat_activity
where backend_type = 'client backend';
Run it as a role that can see every session’s details: PostgreSQL’s docs say ordinary users see all the information only about their own sessions, while superusers and roles with privileges of the built-in role pg_read_all_stats see it for all sessions. PostgreSQL also keeps some slots for superusers, so ordinary connections are refused before the count reaches max_connections: one more reason to alert well below the line.
If the provider has no alert on this number, a scheduled check that runs the query every few minutes and posts to your alert channel when it crosses the line is enough. When the alert fires, the causes and repairs are in every connection pool exhausted error and its fix.
Queue depth alerts on a managed queue: the SQS metrics that matter
The SQS metrics in CloudWatch that carry the backlog alarm are three: ApproximateNumberOfMessagesVisible on the queue, ApproximateAgeOfOldestMessage, and ApproximateNumberOfMessagesVisible on the dead-letter queue, the metric AWS recommends for a DLQ. The sent, deleted and empty-receive counts are dashboard context, not alarms.
These AWS SQS metrics are published to CloudWatch under the AWS/SQS namespace, and AWS’s list of SQS CloudWatch metrics defines each one; its monitoring tips say to set CloudWatch alarms based on ApproximateNumberOfMessagesVisible “to catch backlog growth.” The table keeps AWS’s definitions and adds the alarm I’d start with.
| Metric | What AWS says it means | Alarm to set (my starting value) |
|---|---|---|
ApproximateNumberOfMessagesVisible on the queue | ”The number of messages currently available for retrieval and processing” | Visible messages rising for about an hour |
ApproximateAgeOfOldestMessage | ”The age of the oldest unprocessed message in the queue”, in seconds | Age above the longest a customer should wait, for two periods in a row |
ApproximateNumberOfMessagesVisible on the dead-letter queue | The metric AWS recommends “to monitor the state of a DLQ” | Any value above zero |
AWS attaches caveats that matter before you trust the alarm. Many values are approximate. On a standard queue, a message received three or more times and not deleted moves to the back of the queue, and poison-pill messages, received again and again but never deleted, are excluded from the age metric until successfully processed. A message moved to a DLQ after exceeding maxReceiveCount has its age reset. So a message that keeps failing can drop out of the age alarm; with a dead-letter queue set up, the DLQ alarm is where it shows up.
The rest are context for a dashboard. ApproximateNumberOfMessagesNotVisible counts in-flight messages received but not yet deleted or expired, NumberOfMessagesSent leaves out messages moved to a DLQ automatically, and NumberOfMessagesDeleted, NumberOfMessagesReceived and NumberOfEmptyReceives describe consumer activity.
The setup is a CloudWatch alarm on the metric with an SNS action. AWS’s steps for an SQS CloudWatch alarm go from Alarms and Create Alarm to Browse Metrics, SQS, then the queue name and metric, and end with an existing SNS topic or a new list of email addresses. The email addresses on a new SNS topic must be verified first, and an alarm that changes state before then delivers nothing.
If the queue is a database table or a Redis list, the same two numbers come from one query: how many jobs are pending, and how old the oldest pending one is.
How to set up API spend alerts
API spend alerts are threshold alerts with money as the metric. My starting set is alerts at about 50, 80 and 100 percent of the monthly budget, plus a daily alert at about 3 times a normal day, which is the one that catches a runaway loop before the month’s budget is gone.
| Alert | Starting value (my working rule) | What it catches |
|---|---|---|
| Monthly budget | Budget used reaches about 50, then 80, then 100 percent | A slow climb across the month |
| Daily spend | About 3 times a normal day | A loop or an abused endpoint, within a day |
| The app’s own per-user usage meter | An account crosses its own limit | The earliest signal, because it fires per account |
These are alerts on what your app spends on paid APIs, not on a marketing platform’s API quota. Where each alert is set on each provider, and which settings only send an email while others actually stop usage, differs from provider to provider and belongs with how to cap monthly usage on a metered API.
A production alerting checklist, and how to test that each alert threshold fires
A production alerting checklist for a small SaaS has 5 conditions: error rate, latency, connection pressure, queue backlog and spend. Each row records the metric, threshold, window, where it is set and how to fire it. An alert counts only after it has been fired on purpose and reached a person.
Copy the table into a doc or a sheet and fill the last column as you go; the starting thresholds come from the sections above. The last two rows belong to other jobs and stay on the list so nothing falls between them.
| Condition | Metric | Starting threshold | Window | Where it is set | How to fire it | Fired on (date) |
|---|---|---|---|---|---|---|
| Error rate | Failed requests divided by all requests | Warning about 2 percent, page about 5 percent, at least about 10 failures | 5 minutes | Error tracker or host metrics | Temporary failing route | |
| Latency | p95 response time on key routes | About twice normal, or about 2 seconds | 10 minutes | Error tracker spans, host metrics or an outside check | Temporary slow route | |
| Connection pressure | Connections in use divided by max_connections | About 80 percent | 5 minutes | Database provider page or a scheduled query | Lower the threshold below the current value | |
| Queue backlog | Visible messages, age of the oldest, DLQ visible messages | Rising for about an hour; age past the customer’s wait for two periods; DLQ above zero | Per alarm period | CloudWatch alarm or a scheduled query | Stop the staging worker, or lower the age threshold | |
| Spend | Spend against budget | Monthly budget at about 50, 80, 100 percent; a day at about 3 times normal | Month; day | Provider billing alerts and the app’s usage meter | Set the budget alert just under month-to-date | |
| Outside uptime check | Up or down, from outside | Set in the uptime monitor | The monitor’s interval | Uptime monitor | The monitor’s own test | |
| Owner and route | Who receives each alert | A named person for every alert | Always | Alert routing | Read the message as received |
My own drill for filling the last column is a working rule, not something a tool requires: run it on staging where you can, and in a quiet hour where you can’t, and lower thresholds rather than break production.
- 01 Errors: add a temporary route that throws an unhandled error, so the tracker records an event and the response is a 500. Put it behind authorization, call it in a loop until the share crosses the line, then delete it.
- 02 Latency: add a temporary route that sleeps longer than the threshold, under the same rules. A rule scoped to named routes will not see it, so for that rule lower its threshold below the current value instead.
- 03 Connections: lower the threshold to just below the current value, wait out the window, then restore it. Never exhaust a production pool on purpose.
- 04 Queue: on staging, stop the worker and send messages until the alarm fires, or lower the age threshold on the real queue.
- 05 Spend: set the budget alert just under the month-to-date figure, wait out the provider's reporting delay before calling it missed, then restore it.
- 06 For each condition, record the time it began, the time the alert arrived, who received it, and that it cleared.
The failing route can be this small in an Express-style app, where requireAdmin stands for your own admin check:
// Temporary alert drill route. Delete it after the drill.
app.get('/drill/fail', requireAdmin, () => {
throw new Error('alert drill: intentional failure');
});
For step 5, the provider’s delay can be long: AWS Budgets information “is updated up to three times a day”, and updates typically come 8 to 12 hours after the previous one. The failures worth finding are dull ones: the rule sits on the wrong environment, the alert goes to a channel nobody reads, or the threshold never clears. Keep the filled table with its dates and each alert message as it arrived, and run the drill again after any change of tool or route. This drill tests thresholds; the error logging article has its own drill for the logging pipeline.
In the Production Hardening Sprint, deliverable 8.3 is verified this way: “Trigger each configured condition and record its alert threshold and behavior.”
While an alert is open, customers only learn about it if you tell them, which is what a status page is for.
Where the sprint does this
In the sprint, deliverable 8.3, Operational threshold alerts, is where we “Alert on error spikes, latency, connection pressure, and queue backlog”, verified by the line quoted under the checklist above. The result goes into the production readiness report, deliverable 13.1, which accounts for all 123 scope items, keeps failures visible until resolved and explains genuine non-applicable items. Hosting, paid tools and API usage remain in your accounts. Deliverable 8.3 sits in the monitoring area of the published scope.
Common questions about alerts and queue metrics
What are custom metrics?
Custom metrics are numbers your own code publishes, such as jobs waiting or sign-ups per hour, beside the ones a platform publishes for you. In CloudWatch’s terms a metric is “a time-ordered set of data points that are published to CloudWatch”; many AWS services provide metrics at no charge, and publishing your own application metrics is charged. The pending count from a queue kept in a database table is a good first one.
How to get sentry alerts on phone?
Connect a Sentry monitor (the error tracker, not the car feature) to an alert whose action reaches a phone you already carry: email, Slack, or an on-call tool such as PagerDuty, all actions Sentry’s alert docs list. On the free Developer plan that means email, because Sentry’s pricing page lists “Alerts and notifications via integrated tools” from the Team plan up. Opsgenie also appears in Sentry’s list, but Atlassian’s end of sale for it took effect on June 4, 2025 and Opsgenie “will be shut down” on April 5, 2027, so it is a poor choice for a new setup. Who should be woken, and when, is the routing decision.
What does SQS stand for?
SQS stands for Simple Queue Service, AWS’s managed message queue. It publishes its metrics to Amazon CloudWatch under the AWS/SQS namespace, which is where the queue alarms on this page are set.
What are AWS metrics?
AWS metrics are time-ordered sets of data points that AWS services publish to Amazon CloudWatch, each in a namespace such as AWS/SQS. A CloudWatch alarm watches a single metric over a time period and acts on its value relative to a threshold, which is how each AWS alert on this page is built.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase