A five-person SaaS does not need to hire an SRE; it needs the 8 things that hire would set up, done once, and one named person on call. So what is a site reliability engineer? A software engineer who treats the live app as the product: they write code, alerts and runbooks so the app stays up, fails safely and recovers fast.

What is a site reliability engineer: the role in plain words

A site reliability engineer is a software engineer whose product is the running system instead of a feature. The outcome is reliability: the app answers, answers correctly, answers fast enough and comes back when it breaks. The method is writing code, alerts and runbooks so each failure costs less the next time.

That is the whole job in three parts, and for a company of five it is one answer among several to a wider question, hardening SaaS applications for resilience. The first part is who: an engineer who writes software, not an operator who follows a checklist. Google’s book says its SRE teams focus on hiring software engineers “to run our products and to create systems to accomplish the work that would otherwise be performed, often manually, by sysadmins.” The second part is what: reliability, which a customer feels as an app that loads, gives the right answer, does it quickly, and returns after an outage. The third part is how: every failure gets turned into code, an alert or a written step, so the next one is cheaper.

Site reliability engineering is the practice, a site reliability engineer is the person, and SRE means either one; the sentence around it tells you which. In the name, I read “site” as whatever the customer touches: the app, its API, the checkout, the emails it sends.

What separates the job from a traditional operations role is a limit on operations work. Google places a 50% cap on the aggregate “ops” work of all its SREs, meaning tickets, on-call and manual tasks, and its rule of thumb is that an SRE team must spend the remaining 50% of its time on development. This definition of site reliability engineering is my plain-words reading of the introduction to Google’s SRE book, where the cap is set out.

Google site reliability engineering: where the role came from

Google site reliability engineering began when a software engineer was asked to run a production team, and Google documents it in free books online. The organization those books describe assumes many services and dedicated teams. The ideas in them work for a team of five.

The book’s introduction is written by Benjamin Treynor Sloss, and his definition is one line: “SRE is what happens when you ask a software engineer to design an operations team.” He joined Google in 2003 and was tasked with running a production team of seven engineers, having only ever done software engineering. Google’s free SRE books are Site Reliability Engineering, The Site Reliability Workbook, and Building Secure & Reliable Systems, and they are the primary source for the terms used below.

My reading of what Google’s version assumes: many services, a team for each important one, and a separate SRE organization that can push work back when a service is too noisy. The book describes that pushback as “shifting some of the operations burden back to the development team” when SREs spend too little time on development. Its on-call chapter puts the minimum for a single-site team covering a service around the clock, with a primary and a secondary on call, at eight engineers. A company of five cannot staff that, so what survives the move is the set of ideas in the next section, not the organization chart.

The book also describes a gate a service passes before SRE takes it on: a production readiness review is “considered a prerequisite for an SRE team to accept responsibility for managing the production aspects of a service.” A small app can run its own version of the production readiness review before a launch.

Why it matters for a five-person SaaS: the work exists whether or not anyone has the title

Reliability work does not go away when nobody is hired to do it. It turns into whatever the founder does in the middle of the night, with no plan and no alert that says where to start. That is my reading, and the audit numbers below point the same way.

In the 21 third-party apps I audited, 17 had no error tracking or alerting: when a user hits an error, nothing records it. At least 17 of the 21 had no deploy gate, and so did all 5 founder apps: every push ships straight to production with nothing checking it first. Across the same 21 third-party apps, the Reliability & Correctness pillar of my scoring averages 31.4 out of 100. The 21 third-party apps and the 5 founder apps are ones I audited in June and July 2026, a selected set that came to me, not a random sample and not a rate for AI-built apps in general.

The table shows four common failures and who notices first. These are illustrations, not incidents from those audits.

What happenedWho catches it in a company with SREsWho catches it in a team of five
A webhook stopped arriving and nobody knewThe alert on missing events pages whoever is on callThe first customer who emails, unless monitoring and an alert channel exist
A Friday deploy broke checkoutThe deploy gate or staged rollout stops it, and the rollback has been practicedThe first customer who emails, unless a deploy gate and a rehearsed rollback exist
A provider outage took the whole app down with itFailure handling keeps the rest of the app up while on-call watches the providerThe first customer who emails, unless the code degrades politely and someone is named
A bad migration needed a restore nobody had rehearsedThe restore has been drilled, so the team knows how long it takesThe first customer who emails, unless backups are verified and a restore was tried once

The pattern in the right-hand column is my reading: the first customer who emails is the default alarm until one of the eight areas below exists. When it happens anyway, start with what to do first when the app is down. Using AI to build the app was not the mistake, and none of this says your app is unreliable. It says what my audits found and what you can check.

How it works: what a site reliability engineer does, and what to do without one

What site reliability engineers do fits into four parts below: the eight areas the role owns, the three ideas worth borrowing, what a job ad asks for, and the version a small team can run without the hire. The two main tables are here.

The site reliability engineer role: the 8 things it owns

The site reliability engineer role owns 8 areas: monitoring and alerting, incident response, postmortems, reliability targets, release safety, failure handling in the code, capacity and performance, and recovery. Automation runs through all of them, because any chore done twice is a candidate for code.

In Google’s book, the responsibilities of a site reliability engineer team are, in general, “the availability, latency, performance, efficiency, change management, monitoring, emergency response, and capacity planning” of its services. The table below is my description of the site reliability engineer role for a small SaaS: the grouping is mine, the vocabulary comes from the book’s chapters on service level objectives, toil and on-call. Of the eight, the three a small team forgets first, in my reading, are reliability targets, release safety and recovery.

AreaWhat the SRE does in itThe question it answers
Monitoring and alertingDecides what is measured and what pages a personWould you know it broke before a customer tells you?
Incident responseSets severity levels, who is called, and what the first hour looks likeWho does what when it breaks?
PostmortemsWrites a blameless account and the actions that followWhy did it break, and what stops it recurring?
Reliability targetsAgrees an availability or latency target and what happens when it is missedHow reliable is reliable enough?
Release safetyRuns deploy gates, staged rollout, and a rollback that has been triedCan a bad release be caught or undone?
Failure handling in the codeAdds timeouts, retries, degradation and health endpointsDoes one broken provider take everything down?
Capacity and performanceRuns load tests, sets limits, watches costWill it hold at the next launch?
RecoveryKeeps backups, drills restores, plans disaster recoveryCould you get the data back, and how long would it take?

Exactly how a site reliability engineer spends the week follows from the cap: operations work such as tickets and on-call takes at most half, and the rest should go to project work that uses coding skills. The book wants postmortems written for all significant incidents, “regardless of whether or not they paged”, under a blame-free culture. The automation point at the top of this section is my own rule, not a line from the book.

SLOs, error budgets and toil: the 3 ideas worth borrowing at any size

SRE’s 3 portable ideas are the service level objective, the error budget and the cap on toil. An objective is a target for a measured service level. The budget is the unreliability that target allows, and releases can spend it. Toil is repetitive operational work, and cutting it is part of the engineering in the job title.

Google’s chapter on service level objectives separates three words. An SLI is “a carefully defined quantitative measure of some aspect of the level of service that is provided.” An SLO is “a target value or range of values for a service level that is measured by an SLI.” An SLA is “an explicit or implicit contract with your users that includes consequences of meeting (or missing) the SLOs they contain.” Full definitions, with examples, are in the SRE book’s chapter on service level objectives.

The error budget is one minus the availability target. The book’s Embracing Risk chapter describes the policy: as long as there is budget remaining, new releases can be pushed, and if SLO violations use it up, “releases are temporarily halted” while effort goes into making the system more resilient; it also names softer versions, such as slowing releases down or rolling them back when the budget is close to used up.

Toil, in the SRE book’s chapter on toil, is “the kind of work tied to running a production service that tends to be manual, repetitive, automatable, tactical, devoid of enduring value, and that scales linearly as a service grows.” Google’s stated goal is to keep it below 50% of each SRE’s time, because left unchecked it tends to expand and can fill all of it.

IdeaGoogle’s definition in one lineThe five-person version
Service level objectiveA target value or range for a service level measured by an indicatorOne availability target for the one journey that makes money, read monthly from the uptime monitor
Error budgetOne minus the availability target; releases continue while budget remainsAfter a week with a customer-facing outage, the fix ships before the next feature
ToilOperational work that is repetitive, automatable and grows with the serviceA list of the three chores repeated every week, and a date for automating the first

The third column is my working rule, not Google’s. What a given target allows in minutes of downtime a month is worked through in what 99.9 percent uptime actually means for a small app.

Site reliability engineer requirements: what a job ad asks for, and why you are not writing one yet

Site reliability engineer requirements in GitLab’s public job description cover systems thinking, Linux and the shell, configuration management, programming in Shell and Ruby or Go, and tools such as Kubernetes and Terraform, plus an on-call rotation. That profile pays off when operations work would otherwise consume an engineer, which a single app on a managed host rarely reaches.

GitLab’s handbook page is open, and read from the employer’s side it shows what the hire is for. GitLab lists the skills needed for a site reliability engineer who is “responsible for keeping all user-facing services and other GitLab production systems running smoothly,” and the role includes building “monitoring that alerts on symptoms rather than on outages.” That is a large production estate with many moving parts. GitLab frames the skills as inclinations and says a candidate may be a fit with some of them.

My reading of a team of five is different. A single app on a managed host with a managed database already has much of the machine-level work, such as patching servers, done by the provider. What is left is the list in the role table, and most of it is setup plus a small, regular slice of one person’s week.

My working rule for when the hire is getting close is any one of these four signals:

  • More than one paging incident a week, for about a month.
  • A second or third service with its own data store.
  • A contractual uptime commitment with penalties.
  • An on-call load one person cannot carry.

For that last signal, Google’s on-call chapter is the yardstick: it caps on-call at no more than 25% of an SRE’s time and treats 2 incidents per 12-hour shift as the maximum. Salary, certification and career paths are questions for candidates, and this page leaves them out.

Site reliability engineering at a startup: what a five-person SaaS does instead

Site reliability engineering at a startup of five is the role’s 8 areas split in two: the controls set up once, and the small part that needs a person. The standing part is one named person on call with stated hours, a daily look at one alert channel, and a written postmortem after every customer-facing outage.

AreaSet up onceWhat still needs a personHow often
MonitoringAn error tracker, an outside uptime check, a few threshold alerts and one channelReading the channelDaily
IncidentsSeverity levels and a one-page planA named person on call, with stated hours, who gets the alertsEvery week, by rota
PostmortemsA templateWriting oneAfter every customer-facing outage
TargetsOne availability targetReading itMonthly
ReleasesA deploy gate and a rehearsed rollbackUsing themEvery release
Failure handlingTimeouts, retries, degradation and a health endpointReviewing themWhen a provider is added
CapacityOne load testRepeating itBefore a launch or a pricing change
RecoveryVerified backups and one restore drillRepeating the drillAbout quarterly

The “how often” column is my working rule. The monitoring row starts from application monitoring best practices. The incident row needs shared words for severity, beginning with what Sev1 means, and the rest fits on an incident response plan template written once. Routing is the part that turns a channel into a person: Slack alerting and who is on call. The postmortem row rests on knowing what RCA means in software.

Here is how the missing column shows up. A small SaaS has an error tracker and an outside uptime monitor, both posting to one shared team channel, and nobody is named on call. Overnight a provider outage takes the payment step down. Both tools post to the channel, each person who glances at it assumes someone else has it, and the first action comes after a customer writes in the next morning. The one-time setup worked. The missing piece was a named person with stated hours, which is the on-call idea in Google’s on-call chapter: an on-call engineer is available to act on production “within minutes, according to the paging response times agreed to by the team and the business system owners.” The rest is my reading of the lesson: the tools in the “set up once” column did their job, and the column that needs a person was empty.

The standing cost is a small, regular slice of one named person’s week, plus whatever the outages cost. When that is no longer enough, getting outside help is a buying decision of its own: fractional CTO vs agency, or a fixed package. The bar all of this setup aims at is what makes code production ready.

How to check your own app

Reliability without an SRE is checked with 6 tests run on purpose: a test error that must alert, a failing monitored URL, an alert read by a stranger, a rolled-back release, one integration switched off in staging, and a recent backup restored into a copy.

These are checks for you to run, not a record of anything run for this page. Checks four to six run on staging or an isolated copy, never on the live app.

  1. 01 Throw a test exception from a test route in production, a real thrown error rather than a not-found page, and time it. Pass: the event shows in the error tracker with a readable stack trace, and a notification reaches the named person by email or in the channel the alert rule names. Fail: no event, or an event nobody was told about. Evidence: the event link and the minutes it took.
  2. 02 Add a monitor on a test URL that returns a server error, or pause the route, and wait at least one check interval. Pass: a down notification, then a recovery message once it is fixed. Fail: silence, or no recovery message. Evidence: both messages with their times.
  3. 03 Send a test alert and read it as a stranger would. Pass: it says what broke, where, and the first step to take. Fail: a bare error name or a link with no context. Evidence: the alert text.
  4. 04 On staging, roll back a test release. Pass: the app works on the previous version and the data written since is intact. Fail: the rollback breaks the app or loses rows. Evidence: before and after screenshots and a row count.
  5. 05 On staging, switch off one integration, such as payments or email. Pass: the rest of the app keeps working and the affected feature fails politely. Fail: a blank page or a spinner that never ends. Evidence: the journeys you tried and what each showed.
  6. 06 Where your plan includes backups, restore a recent one into an isolated copy and check three records you know. Pass: all three match. Fail: no backup, or a restore you cannot complete. Evidence: the three records and how long the restore took.

On Supabase, daily backups are automatic on the Pro, Team and Enterprise plans, and restoring into a new project is limited to paid plans with physical backups enabled; on the free tier, Supabase recommends regular exports with the CLI’s db dump command and off-site backups. Sentry’s alert rules can send a test notification from the rule itself, which verifies the wiring before you save; Sentry notes that event-based triggers may behave differently, so the thrown error in the first check still matters.

Write down the date, the result and the time each check took. Together the six answer the question an SRE is paid to answer: would you know, and could you get back? Two more need no tool. Name the person on call this week out loud, and find the last written postmortem. If either is missing, that is the first fix.

In the Production Hardening Sprint, deliverable 6.3 is verified this way: send a test error and verify symbolication, environment attribution, and alert delivery; deliverable 8.2 is verified this way: trigger a controlled check failure and verify notification and recovery reporting.

Where the sprint fits

Area 06 of the sprint scope, Error handling & reliability, contains 10 of the 123 deliverables. The ones this page touches are: 6.3, install error tracking such as Sentry with protected source maps, environment labels, and alerts; 6.8, keep unaffected functions usable when an external service is unavailable; 8.2, monitor the production URL and health endpoint with outage alerts; 8.4, route alerts to the designated Slack or email destination and tune thresholds to reduce noise; 7.6, document and rehearse release rollback, including how database changes are handled safely; 4.7, verify automated backups, set retention, and perform a real restore test; and 13.3, document deployment, rollback, key rotation, backup restoration, and the response to each operational alert. Post-handover support is 14 calendar days of fixes for defects in the delivered sprint work, and hosting, paid tools, and API usage remain in your accounts. Every deliverable is listed in the published scope.

Common questions about the SRE role

Are SRE and DevOps the same?

No. Google’s SRE workbook calls DevOps “a broad set of principles about whole-lifecycle collaboration between operations and product development,” and SRE “a job role, a set of practices” plus “some beliefs that animate those practices.” Its shorthand is “class SRE implements interface DevOps”: in the workbook’s reading, SRE implements some of the philosophy that DevOps describes.

The workbook also says SRE believes in the same things as DevOps but for slightly different reasons. The other half of the comparison starts from DevOps in simple terms.

What are the four pillars of SRE?

The SRE book’s contents name no four pillars. They run in five parts, and the part called Principles has seven chapters, from Embracing Risk to Simplicity, with Service Level Objectives and Eliminating Toil among them. For a small team, the eight areas in the role table above are the working list.

Is SRE a stressful job?

Yes, it can be, and the part Google sets limits on is on-call load: its on-call chapter allows no more than 25% of an SRE’s time on call and treats 2 incidents per 12-hour shift as the maximum. The same chapter notes that night shifts have detrimental effects on people’s health.

A founder carrying the pager alone has none of those limits unless they set them. My reading is that the same fixes apply at any size: stated hours, a second person who can take over, and fewer pages through better alerts.

Is AI replacing SRE?

No, not the part that decides outcomes. AI tools can shorten toil such as searching logs or drafting a runbook step, but deciding what to measure, what should page a person, and whether to roll back stays with a person who owns the result (my reading). An assistant sees only what you describe to it or connect it to, and it does not carry the pager.