The night the checkout stops taking cards is a Sev1: severity one, the incident that cannot wait. Agree beforehand what it obliges: someone is woken, customers are told, and everything else stops until it is fixed. A five-person team needs 3 levels, each tied to who acts, how fast, and what customers hear.
Sev1 meaning: the core job of the app is down, and it cannot wait
Sev1 means a severity-one incident: the product’s main job is down for most customers with no workaround, or customer data or money is at risk for any customer. Sev is short for severity. Ladders differ between teams, and PagerDuty’s public one runs from SEV-1, the most urgent, down to SEV-5 for cosmetic issues.
Deciding how serious each kind of failure is, before anything fails, is one of the cheapest parts of hardening SaaS applications for resilience. The whole ladder fits on one page.
In tech, the meaning of “sev” is plain: a sev is an incident that has been given a severity level, and Sev1 is the most serious one on the ladder. PagerDuty says its own descriptions “have been changed from the PagerDuty internal definitions to be more generic”, and it encourages teams to make theirs “very specific, usually referring to a % of users/accounts affected.” That advice is why the levels below are written as tests, not adjectives. A guide can also count from SEV0: Xurrent’s is titled “SEV0-SEV5 explained”, though its text concentrates on SEV1 through SEV3.
One published ladder that ties every level to a response time is a cloud provider’s support scale. AWS Support’s case severity levels run from “General guidance”, with a 24-hour first response, to “Business-critical system down”, with less than 30 minutes, less than 15 minutes or 5 minutes depending on the plan. AWS words those times as an effort, not a promise: it “makes every reasonable effort to respond to your initial request within the indicated timeframe.” Every level is listed under AWS Business Support+, AWS Enterprise Support or AWS Unified Operations, and “If you have Basic Support, you can’t create a technical support case,” so a small team on the Basic plan meets this ladder only in AWS’s docs (checked 30 September 2026).
Security teams start from a different place. NIST’s computer security incident handling guide, NIST SP 800-61 Rev. 2, published in August 2012, says handling “should be prioritized based on the relevant factors” and names three: functional impact, information impact and recoverability. NIST withdrew it when its April 2025 revision superseded it.
Two other things share the name and are not covered here: the Seversky SEV-3, an American three-seat amphibian aircraft that first flew in 1933, and AMD’s SEV, Secure Encrypted Virtualization, a feature of AMD processors. A level on a ladder is only useful once everyone knows what it makes them do.
What it means in practice for a small SaaS: three levels, and what each one obliges
Three severity levels are enough for a five-person team. Sev1: the main job is down, or data or money is at risk, so someone is woken and customers are told. At Sev2, something important is broken for some customers, handled the same day. Sev3: a real problem that waits for working hours and the normal queue.
Picture a five-person SaaS with one engineer and no on-call rota. Nobody’s job title says “incident commander”, so the table has to say who acts. Every time in it is my working rule for a team that size, an approximate one, and none of it is an industry standard or an SLA.
| Level | The test | Examples | Who acts, and how fast | What customers are told, and where | Postmortem owed |
|---|---|---|---|---|---|
| Sev1 | The main job is down for most customers with no workaround, or any customer’s data or money is at risk | Login fails for everyone; checkout stops charging or charges twice; one customer can see another’s data | Whoever holds the phone is woken at any hour and acknowledges within about 15 minutes; all other work stops | One short message inside those same first 15 minutes, then updates at the times you promised, on your status channel and to anyone affected | Yes |
| Sev2 | An important feature is broken or the app is badly degraded for some customers, or for everyone with a workaround | Exports fail; emails are delayed by hours; the app is slow enough that people give up | Same-day response in waking hours | The affected customers, directly | Only if it lasted more than a few hours |
| Sev3 | A real problem that can wait for working hours | A minor feature is broken; a cosmetic fault; one customer’s edge case with a workaround | A ticket, fixed in the normal flow of work | A reply to the customer who reported it | No |
The Sev1 clock lines up with what to do first when your app is down, which puts confirming the outage, capturing the logs and telling customers in the first quarter of an hour; a severity ladder should never set a slower clock than that. A partly broken app with people paying for it is usually the Sev2 case, and my app is broken and I have paying customers covers what to switch off and how to word the messages; once money is moving wrongly, the same case is a Sev1 whatever the count. What to publish during an outage, and whether you need a public page at all, belongs to a status page for a small SaaS.
A postmortem owed means a written review of the cause, not only of the fix, and what RCA means in software is the method behind it. Who holds the phone, in a team with no rota and nobody on call by title, is a decision of its own that starts with what a site reliability engineer is, and what a small team needs instead. Once the three levels are agreed, they go into the severity field of an incident response plan template. A security incident is always handled under that plan, whatever level it first gets.
A quiet failure can matter more than a loud one. In an AI coding workspace I audited in June and July 2026, every database call went through one helper that logged and returned a harmless fallback on any error. The only caller of the chat-save function ignored its result, so a failing save looked like success and a user’s work could vanish. My reading of it: the error went to a log line while the screen said the save worked, so once someone does notice, the level is set by the lost work, not by how loud the error was.
Sev 3, and everything below it
Sev 3 is, by my working rule, the lowest level a small team needs: a minor feature broken, a cosmetic fault, or a single customer’s edge case with a workaround. It gets a ticket, a reply and a fix in normal working hours. Below that, it is just the backlog.
A background job that failed belongs here too, as long as it can be replayed safely, which is what an idempotency key for safe retries is for. No postmortem is owed at this level. Published ladders go further down, and PagerDuty’s runs to SEV-5, as the definition above shows. Xurrent’s guide points the same way for small teams: “for smaller teams or startups, these extra levels can sometimes create confusion.”
One more working rule of mine runs the other way: a Sev3 that recurs every week is a Sev2 nobody has added up.
How to pick the level in two minutes
An incident’s severity is picked with 4 questions in order: is data or money at risk, is the main job unusable for most customers, is something important broken for some, or none of these. When unsure, choose the higher level, and my working rule is to downgrade later.
The first yes decides:
- 01 Is any customer's data exposed, or is money moving wrongly, right now? Then it is Sev1.
- 02 Is the main job of the product unusable for most customers, with no workaround? Then it is Sev1.
- 03 Is something important broken for some customers, or for everyone but with a workaround? Then it is Sev2.
- 04 If none of these, it is Sev3.
When two levels both look possible, take the higher one. PagerDuty’s guidance says to “treat it as the higher one”, and adds: “During an incident is not the time to discuss or litigate severities, just assume the highest and review during a postmortem.” Google lists “Declare incidents early and often” among the basic principles in the SRE workbook’s incident response chapter. My working rules on top of those: the level can come down later, and whoever sees the problem declares it, in one named channel, without asking anyone’s permission. If the level changes during the incident, write the change down with the time.
That chapter also shows what the missing habit costs. Its first case study, Google’s own account, follows a bug in Google Assistant version 1.88 on Google Home devices, with a timeline that runs from May 22 to June 4, 2017; the support team received “numerous customer phone calls, tweets, and Reddit posts”, and “Despite all the user reports and feedback, the bug wasn’t escalated to a higher priority.” When support finally raised the bug to the highest priority on June 4, “The team did not declare an incident” and kept troubleshooting through the bug tracker. Google’s write-up says “early escalation would have produced a quicker, more organized response, and a better outcome.” The lesson I take from it is that a raised ticket priority is not a declared incident. When reports pile up and nobody has named a level, declare one now, and pick the higher.
What it is not: a priority, a P0, or a verdict on anyone
A severity level gets mixed up with three other things: a priority, a log level and a judgment on whoever caused the problem. The priority mix-up is the common one, so it has its own section below.
It is not the severity in your logs. error and fatal are labels on single log lines, and one Sev1 can produce thousands of them or none at all. Those labels are set by the code, line by line, while an incident’s level is set by a person looking at what customers can and cannot do. Logs severity levels in production are a separate ladder with a separate job.
It is not blame either. The level describes the impact on customers, never the size of someone’s mistake, and the review that follows a Sev1, the RCA mentioned above, is blameless.
P0 incident and P1 to P4: severity versus priority
A P0 incident sits at the top of a priority scale that runs from P0 to P4. Google’s Issue Tracker docs, as a common way of prioritizing, give P0 as an issue that needs to be addressed immediately and with as many resources as is required. Priority is the order work gets done; severity is how bad the impact is.
The last sentence there is my reading, not Google’s wording. Google Issue Tracker’s priority definitions come after a note that “Teams generally have different criteria for how importance of an issue is determined.” The rest of its P0 line says such an issue “causes a full outage or makes a critical function of the product to be unavailable for everyone, without any known workaround.” That is close to the Sev1 test on this page, so a small team can run one ladder and call its top level either name, as long as each level has obligations.
The two scales can split lower down. A typo on the pricing page is low severity and high priority, because every visitor sees it. A rare crash in an admin screen that only you use is the reverse. The “who sets it” column below is my reading of how a small team works, not either source’s.
| Scale | The ladder | What it measures | Who sets it |
|---|---|---|---|
| Severity | SEV-1 (most urgent) to SEV-5 (cosmetic issues) on PagerDuty’s public ladder | How badly customers are hit right now | Whoever declares the incident, from the customer impact |
| Priority | P0 (“addressed immediately”) to P4 (“addressed eventually”) in Google Issue Tracker | The importance of an issue, which sets how soon the work is done | Whoever orders the work queue, usually a founder |
If a customer’s contract uses P1 to P4 with response times, those contract terms override anything on this page, and the contract language is the thing to read.
How to check that your levels work
Severity levels are tested with a tabletop drill, about 30 minutes by my working rule: each person rates three incidents separately and the answers are compared, a real test alert is followed to its destination on the Sev1 phone in night mode, and a test incident is published to the status channel.
A ladder nobody has tried fails on the night it is needed. Run the drill with everyone who might hold the phone, in the room or on the call. Each check below is one the team can fail:
- 01 Read three past or invented incidents aloud and have each person write down a level without conferring. If the answers differ, sharpen the wording of the tests in the table.
- 02 Send a real test alert, through the same tool and channel that would page someone at night, to whoever holds the Sev1 phone with their sleep or night mode on, and confirm it makes a sound. Do the same for the second name, and if there is no second name, choose one now.
- 03 Look at where that alert landed and what it says: does it name what broke, carry enough context to start, and say what to do first?
- 04 Publish a test incident to the status channel with subscriber notifications off, or on a test page, check it loads from a host that is not the app's own, then delete it.
An alert that lands in email or in a muted chat channel has not woken anyone, however fast it arrived. Write down the date of the drill and what it changed, and run it again whenever the team or the hosting changes.
In the Production Hardening Sprint, deliverable 8.4, Actionable alert routing, is verified this way: “Send test alerts and verify their destination, context, and response instructions.” Deliverable 8.7, Public status page, is verified this way: “Publish a test incident and verify the page remains reachable separately from the app.”
How it shows up in a hardening sprint
Four of its deliverables sit nearest to what a Sev1 obliges: alert routing, the status page, an outage rehearsal and the runbooks. Alert routing (8.4) and the status page (8.7) are verified as quoted in the drill section. Deliverable 6.10, Staging outage simulation, is verified by recording “outage scenarios, queue behavior, customer messages, and recovery results.” Deliverable 13.3, Operating runbooks, is verified when we “Walk through the runbooks against the delivered configuration and reference the rehearsal evidence.” After handover, the included cover is 14 calendar days of fixes for defects in the delivered sprint work and 30 calendar days of async questions about the handover and architecture. Each of those lines is in the published scope.
Common questions about incident severity
Is sev 1 higher than sev 2?
Yes. Sev 1 is the more serious level, and PagerDuty’s public ladder states the rule: severities are classified “with the lower numbered severities being more urgent.” Some people hear a bigger number as a bigger problem, so say the word with the number on a call: “sev one”, not “level one”.
What’s the highest level of severity?
The highest level is Sev 1 on a ladder like PagerDuty’s public one, where SEV-1 is the most urgent of five and SEV-5 is the least. Xurrent’s guide counts from SEV0 in its title, “SEV0-SEV5 explained”, but its text gives no SEV0 definition. My working rule for a five-person team is to skip a level above Sev1: keep the top level for the night the main job stops or data or money is at risk.
What is P0, P1, P2, P3, P4 level priority?
P0 to P4 is a priority scale, and Google’s Issue Tracker gives it as “a common way of prioritizing issues”: P0 is addressed “immediately”, P1 “quickly”, P2 “on a reasonable timescale”, P3 “when able” and P4 “eventually”. Priority mostly says how soon, but Google’s rows also describe impact, such as a full outage for P0, which is why a P0 and a Sev1 can describe the same night.
What does SEV 1 mean at Amazon?
At Amazon, SEV1 labels a high severity incident: an Amazon job posting names “high severity incidents (SEV1/SEV2)” and defines neither level. How Amazon defines it internally is not published on Amazon’s or AWS’s sites (searched 30 September 2026), so anything more specific you read is secondhand. What AWS does publish is its support case scale for customers, with names instead of numbers, from “General guidance” to “Business-critical system down”, a first-response time for each, and every level listed under the Business Support+, Enterprise Support or Unified Operations plans. One AWS product line does number its support severity levels: AWS Elemental’s run from “Severity Level 1: Urgent Issue - Production Blocking” to Severity Level 4.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase