No error tracking, no alerting, so a user’s error goes unrecorded: I found that gap in 17 of 21 third-party codebases from my audit round of June and July 2026. Slack alerting is half the fix. Each alert must land in a channel someone owns, say what broke and for whom, and link the first step.
What actionable alert routing is
Alert routing that gets acted on means each alert goes to a destination with an owner and carries 5 fields: what broke, where, since when and how often, who is affected, and the first step with a link. An alert without the last field gets read and left.
The 21 apps behind that opening number were third-party apps: 11 public ones I audited across all 12 pillars, and a held-out set of 10 I audited blind. They are a selected set of audited apps, not a random sample, and not a rate for AI-built apps in general.
Routing is one part of logging and monitoring for a small SaaS: the decision of where each alert goes and who owns it once it fires. Which signals deserve watching at all is the subject of application monitoring best practices. In my reading, an alert can be acted on when three things hold: the destination is right, the alert carries context, and it carries response instructions. The five fields below are my list, each with a bad and a good version.
| Field | Bad version | Good version |
|---|---|---|
| What broke | A raw stack trace from the checkout route | Checkout is failing: payment confirmation returns an error |
| Where | No environment named | Production, the payments service |
| Since when and how often | One message for every failed request | One message with the time of the first failure and a count over a fixed window |
| Who is affected | Not stated | Every customer on a paid plan who tries to upgrade |
| The first step | Check the logs | A link to the saved log query for checkout errors and a link to the checkout runbook page |
This table is the routing version of a rule that log-based alerts already follow: alert on a decision someone has to make, not on every log line.
What goes wrong without it
The table lists four failures, why each one happens and the fix for it on this page. The middle column is my reading.
| What happens | Why | The fix on this page |
|---|---|---|
| No alert at all | Nothing watches the app, so the customer tells you | An alert for each condition that needs a person, routed to an owned channel |
| The alert goes to one person’s email | That person is asleep, on a plane or gone from the company | An owned channel plus a second route for pages |
| Too many alerts and everyone ignores them | Every exception posts a message, the channel gets muted, and the one alert that mattered scrolls past | The weekly noise pass |
| Every vendor’s status feed posts into your alerts channel | The host, the database, the payment and the email providers are all subscribed for every component | A separate low-priority channel, subscribed only to the components the app uses where the page allows it |
The first row is the opening finding from the same audits. Drowning in SaaS status alerts is the last row: every vendor incident lands in the channel your own monitoring uses, and a real outage of yours reads like one more vendor notice. Vendors’ status pages differ in what they let you narrow. Vercel’s status page, for example, offers email, SMS, Slack, webhook and RSS or Atom subscriptions. Status pages built on Atlassian’s Statuspage can show subscribers a component selection screen, a feature only available on Statuspage’s Business plans and higher, so whether you get that choice depends on the vendor’s plan.
The second row is the quietest. Picture a founder whose uptime monitor was set up by a freelancer, with the outage alert going to the freelancer’s own email address. The freelancer has moved on; the app goes down, the monitor notices and sends its alert on time, into an inbox nobody at the company reads. An alert is only useful if the responsible person sees and understands it, and the lesson I take from this case is that an alert sent to a person instead of a role stops working the day that person moves on: route it to an owned channel, then send a test alert. When the app is already down, the order of the first hour is in what to do first when your app is down.
Alert fatigue: what to do when nobody reads the alerts any more
Alert fatigue is what happens when a team receives more alerts than it can act on and starts ignoring all of them, including the real ones. The test is a week’s count: how many alerts fired, and how many of them led anyone to do anything.
Hospitals know the same effect as alarm fatigue: clinicians become desensitized to safety alerts. In a Slack channel that posts every exception, it shows up as a mute button: people silence the channel, and from then on the real alert arrives to nobody.
To resolve alert fatigue, my working rule is to mute nothing, to delete or fix instead, and to start from zero alerts that page, adding back only the ones where the on-call person would do something different because the alert fired. Which conditions deserve an alert in the first place is covered under alert on decisions, not every line.
Slack alerting: one owned channel, a second route, and a person on call
Slack alerting for a small team is 3 tiers and 2 channels. Pages go to an owned alerts channel and a second route that reaches a phone. Tickets go to the same channel for working hours. Everything else stays on the dashboard. Vendor status feeds get their own channel.
Each Slack alert belongs to exactly one tier, and the tier decides where it goes, who must respond and how fast. The table is my working rule, and its last column gives rough ranges, not a standard.
| Tier | Example | Where it goes | Who must respond | How fast |
|---|---|---|---|---|
| Page | Customers cannot log in, pay or load the app | The alerts channel, plus a second route that reaches a phone | The on-call person, then the backup | Within minutes |
| Ticket | A job failed its last retry; the error rate is above normal | The alerts channel | Whoever is on call, in working hours | The same working day |
| Log | Everything else | No notification; it shows on the dashboard | Nobody is notified | No clock |
The log tier is what keeps the other two quiet: a condition worth seeing but not worth interrupting anyone for belongs on a dashboard, and how to build an ops dashboard is a separate job from routing. To set up Slack alerts this way, work through the five sections below in order: the plumbing, the channel settings, the second route, the rota, and the weekly pass.
How to send alerts to Slack: incoming webhooks and app integrations
An alert reaches Slack in one of two ways: the monitoring tool’s own Slack integration, or an incoming webhook, a URL that accepts a JSON message for one channel. Slack’s docs say the webhook URL contains a secret, so it belongs in a server environment variable.
Start with the tool’s own integration where it has one, because it needs no code. In Sentry, you install Slack from Settings, Integrations, Slack, then select Slack as an action on an alert and name the workspace and the channel or user to notify; Sentry’s Slack integration docs add that you can send a test notification to double check. That route has a plan condition: Sentry’s pricing table lists alerts and notifications via integrated tools, and third-party integrations, on the Team plan and above, while the Developer plan has alerts and notifications via email and is limited to one user.
For a tool with no Slack integration, or your own code, the incoming webhook is how to create alerts in Slack. These are the steps as Slack’s incoming webhooks docs give them:
- 01 Create an app in Slack's app dashboard and pick the workspace.
- 02 In the app settings, open Incoming Webhooks and switch Activate Incoming Webhooks on.
- 03 Click Add New Webhook to Workspace.
- 04 Pick the channel the app will post to, then select Authorize.
- 05 Copy the URL listed under Webhook URLs for Your Workspace into a server environment variable.
Then POST a JSON body with a text field to that URL. Keep the URL out of the command itself:
curl -X POST -H 'Content-type: application/json' \
--data '{"text":"Checkout failing in production for paid customers. Runbook: <your runbook link>"}' \
"$SLACK_WEBHOOK_URL"
Slack’s docs are blunt about the URL: “Your webhook URL contains a secret. Don’t share it online, including via public version control repositories. Slack actively searches out and revokes leaked secrets.” That is why it lives in an environment variable on the server and never in client code or a repo, and why a leaked one gets the same treatment as a leaked key: how to rotate API keys safely applies to it too. A webhook is specific to a single channel, and the channel cannot be overridden from the message, so one webhook serves one channel.
Create one channel for pages and tickets, #alerts, and one for vendor status feeds, #vendor-status, and keep people’s conversation out of both. One limit applies on a free workspace: “On the free version of Slack, you’ll be limited to 10 third-party or custom app installations.” The webhook’s app and each tool’s Slack app count as installations in my reading, so a free workspace near that limit can send several tools through one webhook app.
Slack notification settings for an alerts channel
Slack’s notification settings decide whether an alert in the channel ever reaches a person, and only a few of them matter for an alerts channel. Each on-call person sets the alerts channel to notify on all new posts, on desktop through the Notifications icon at the top of the channel, and separately on mobile under the channel’s Settings and Details, because Slack lets the phone differ from the desktop. On mobile, set notifications to arrive as soon as a message is sent rather than after a delay. Then check the notification schedule: outside the hours you set, Slack pauses notifications.
Mentions do not get around a pause. Slack’s help page on notifying a channel says @channel, @here and @everyone “won’t notify people when their notifications are paused, or when they’re used in threads”, and owners and admins with permission can restrict who uses them. My conclusion: a mention in an automated message cannot wake someone whose notifications are paused, so Slack alone is not a pager, which is why the page tier needs a second route.
Email, SMS and the second route
The page tier always has two routes, and at least one of them does not depend on Slack, on the app, or on the provider that is failing. By my working rule, an uptime monitor’s SMS, call or mobile push option is the usual second route.
The second route exists for two cases: Slack is muted or down, or the person is away from a screen. It is whatever reaches a phone and makes a sound. For the uptime check, the uptime guide linked in the verify section already says to pick a channel that actually interrupts, and records that UptimeRobot’s free plan sends email while SMS and voice-call credits are bought separately, so a free push app is usually the cheapest fix there. For every other alert tool, use its SMS, phone call or mobile push option if it has one. Once the team is past about three people, my working rule is that a paging tool such as PagerDuty takes over this job. Opsgenie is not one to start on: Atlassian’s end-of-sale date for it was June 4, 2025, existing customers can use it until April 5, 2027, and then it will be shut down.
Email still has a job: it is the right place for the ticket tier’s daily digest, read once in the morning.
Who is on call when the team is two people
An on-call alerting checklist for a team of two names one person per week and a backup, both with tested logins to the host, database, payment provider and registrar, and makes sure a page reaches the backup too, after a stated number of unacknowledged minutes where the tool can escalate.
This is my checklist for a team of one to three:
- One named person is on call each week, and the name is written where the whole team can see it.
- A named backup covers the same week.
- An unacknowledged page reaches the backup after about ten to fifteen minutes where the alerting tool can escalate; where it cannot, both get the page at once.
- The on-call person has working logins to the host, the database, the payment and email providers and the domain registrar, tested this month.
- The on-call person knows where the runbooks are.
- Acknowledging means a reaction or a reply on the alert in the channel, so the rest of the team knows someone has it.
- Handover happens on a fixed day each week.
- A solo founder has written down who gets the page during a flight, even when the answer is that the status page tells customers the team is away.
The runbooks the alert links to follow a runbook template for a small SaaS. For the status page in the last item, what is a status page for a small SaaS covers what to post during an outage. Once a page turns out to be a security incident, the steps move to an incident response plan template.
Tune alert thresholds to reduce noise: the weekly pass
Alert tuning is a weekly pass over every alert that fired, giving each one a ruling: keep, raise the threshold or add a duration, group, downgrade, or delete. By my working rule, an alert nobody acted on three weeks running does not survive the fourth.
Once a week, list every alert that fired and give each one of these rulings. This is my routine:
| What the alert did last week | The ruling |
|---|---|
| Someone acted on it | Keep |
| Fired on a single spike that cleared by itself | Raise the threshold, or add a duration, such as above the line for five minutes |
| Fired many times for one problem | Group into one message with a count |
| Was real, but could have waited | Downgrade from page to ticket, or from ticket to log |
| Nobody acted on it | Delete |
The threshold numbers themselves belong to how to alert on error rate spikes; this pass only decides whether each alert earns its place. Grouping log alerts by error type and release follows the error-logging guide linked under alert fatigue above. One more rule of mine: a missing alert found after an incident is added the same day. Record the number of alerts per week as well. It should fall over the first few passes and then hold.
MTTD, MTTA and MTTR: the clocks an alert starts
MTTD, mean time to detect, runs from a failure to its alert; MTTA, mean time to acknowledge, from the alert to a person taking it; MTTR can mean respond, repair, recover or resolve, so say which. Mean time to respond runs from the alert until service works again; mean time to recovery runs from the failure itself.
Atlassian’s incident metrics guide says MTTR “potentially represents four different measurements” and that “The R can stand for repair, recovery, respond, or resolve”. MTTD, mean time to detect, is not defined there; I use it in its common sense, the time from a failure starting to the alert firing. The table puts all of them side by side, which also answers what MTTA vs MTTR means for a small team.
| Acronym | Clock starts | Clock stops | What shortens it on this page |
|---|---|---|---|
| MTTD, mean time to detect | The failure begins | The alert fires | Monitoring the right signals |
| MTTA, mean time to acknowledge | An alert is triggered | Work begins on the issue | Routing, the second route and the rota |
| MTTR, mean time to respond | You are first alerted | The product or service is fully functional again | The first-step link in every alert |
| MTTR, mean time to recovery | The system or product fails | It is fully operational again | Everything above, together |
| MTTC, mean time to contain | The incident is detected | The threat is contained | Not this page: incident response |
The MTTD row is common usage and the MTTC row is the security teams’ meaning, not Atlassian’s. In Atlassian’s words, mean time to respond “does not include any lag time in your alert system”, while mean time to recovery “includes the full time of the outage”. So mean time to recover, read that way, includes detection, and that is the clock a customer feels.
Google’s SRE workbook, in its chapter on alerting on SLOs, calls the detection part detection time: “How long it takes to send notifications in various conditions.” DORA’s metrics guide has a narrower recovery metric, failed deployment recovery time: “The time it takes to recover from a deployment that fails and requires immediate intervention.” Mean time to contain comes from security work, and NIST SP 800-61 Rev. 3 words containment as “preventing the expansion of an incident”.
For a small team, my working rule is to write down four timestamps per incident: when it started, when the alert fired, when someone acknowledged it and when it was fixed. MTTD, MTTA and both kinds of MTTR fall out of those four once there are a few incidents to average. Benchmarks are left out on purpose; the last answer in the FAQ says why.
How to verify it
Alert routing is verified by sending a test alert down every route and checking 3 things: it reached the right destination, it carried context, and it carried response instructions. Time the acknowledgement; that is your first MTTA figure.
To test alert routing to on call, run these checks with the real on-call person, not with yourself watching the channel. Each check says what you should see; a check you cannot pass is the finding.
- 01 For each tool that alerts, use its send test alert button, or trigger a harmless real condition such as a test exception in an environment the alert rule covers, or an uptime check pointed at a test endpoint that returns an error. You see the message in its own channel, after the tool's check interval.
- 02 Read the message as a stranger would. You see all five fields, the link opens the right dashboard or log query, and the runbook link works.
- 03 Close the on-call person's laptop and send a test inside their notification schedule. You see the Slack notification on their phone.
- 04 Fire a page-tier test. You see the SMS, call or push on the phone; where the tool escalates, you also see the backup's notification after the stated minutes with no acknowledgement.
- 05 Write down when the alert fired and when someone acknowledged it. You have two timestamps, which give your first MTTA figure.
- 06 Take the on-call person off the rota for one test. You see the backup get the page.
- 07 Break the webhook on purpose in staging with a wrong URL. You see the failure in the sending tool's log, or you write the missing failure down as a gap.
For the last check, Slack’s webhook docs list error responses such as invalid_token, no_service and channel_not_found, so a tool that logs its webhook responses can show one of them when delivery fails. Keep the evidence: a screenshot of each test alert in its destination, the two timestamps per route, and the rota as it stood on the day. The uptime monitor’s own drill is in test the monitor, then keep it current.
On the Production Hardening Sprint, deliverable 8.4 is verified this way: send test alerts and verify their destination, context, and response instructions. For the uptime tool, deliverable 8.2 is verified this way: trigger a controlled check failure and verify notification and recovery reporting.
Where the sprint does this
Deliverable 8.4 of the Production Hardening Sprint, Actionable alert routing, routes alerts to the designated Slack or email destination and tunes thresholds to reduce noise; its check is the one in the verify section above. The threshold alerts themselves are deliverable 8.3, which alerts on error spikes, latency, connection pressure, and queue backlog, and the uptime check is deliverable 8.2. All of them are recorded in the production readiness report, deliverable 13.1, which accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. The fee covers our engineering work: hosting, paid tools and API usage remain in your accounts, and we explain any required third-party costs before enabling them. Each of these deliverables is listed in the published scope.
Common questions about alerts in Slack and response times
How do I alert everyone in a Slack channel?
Put @channel in the message to notify all members of the channel, or @here to notify only the active members. A message sent by an app, such as a webhook alert, works differently: Slack’s help says that for a Slack app that can notify members, the message from your app’s bot user must contain <!channel> or <!everyone>. Neither form gets past paused notifications, as the settings section above explains.
Does @everyone work on Slack?
Yes, but only in the general channel, which all members except guests are added to automatically; there it notifies every person. Owners and admins with permission can choose to restrict who may use it. For an alerts channel, use @channel.
How can I set up Slack notifications to send me an email?
Turn on email notifications in your Slack notification preferences. When you’re not active in Slack, it can email you about mentions, DMs and replies to threads you follow, bundled once every 15 minutes or once an hour. That delay, and the fact that it covers mentions and DMs rather than every message in a channel, is why it does not work as the second route for a page.
Why is Slack not notifying me about alerts?
Usually because your notifications are paused or the channel is set to mentions only. Outside the hours in your notification schedule, Slack pauses notifications, and @channel or @here in an alert does not get past a pause. Set the alerts channel to All new posts on desktop and again on mobile, since the phone keeps its own setting, and have mobile notifications arrive as soon as a message is sent rather than after a delay.
How is MTTR calculated?
Add up the time from the start point to the stop point for every incident in a period, then divide by the number of incidents. The start and stop points depend on which R you mean: mean time to respond counts from the alert until the service works again, and mean time to recovery counts all the downtime, as the clocks table above sets out.
What is a good MTTR value?
There is no universal good value. DORA’s metrics guide defines failed deployment recovery time but publishes no performance bands for it. For a small team, the useful number is your own trend: the same MTTR, measured the same way, falling across your own incidents.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase