Seventeen third-party apps out of the 21 I audited had no error tracking or alerting: when a user hits an error, nothing records it. The RCA meaning software teams use is root cause analysis: working back from what users saw to the conditions that allowed it, then changing those conditions. The postmortem is where it gets written down.
The RCA meaning software teams use: root cause analysis, in plain words
RCA in software means root cause analysis: the work done after service is restored to find why a failure was possible, not only what broke. It separates the symptom users saw, the trigger that set it off, the root causes that allowed it, and the contributing factors that slowed the response.
Those 21 third-party apps are the ones I audited in June and July 2026, a selected set of audited apps, not a random sample and not a rate for AI-built apps in general. Without an error tracker, the first record of a failure may be a customer’s email, and an investigation that starts from an email starts late.
RCA stands for root cause analysis, and in software development it begins once customers can use the app again. It is one part of hardening SaaS applications for resilience: the part that stops the same failure from coming back.
In my framing, an RCA keeps four words apart. The symptom is what users saw, such as a checkout that spins forever or an error page. The trigger is what set it off: a deploy, a traffic spike, an expired key. The root causes are the conditions that let the trigger do damage, such as no timeout on an outside call, no test on the path, no alert on the error. Contributing factors are whatever made it slower to notice or fix. A postmortem that lists only the trigger has described the outage without explaining it.
| Activity | The question it answers | When it happens | What comes out of it |
|---|---|---|---|
| Debugging | What is broken right now? | During the outage, or while the bug is open | A fix, a rollback or a workaround |
| Root cause analysis | Why was this possible at all? | After service is restored | The root causes and the conditions to change |
| Postmortem | What happened, why, and what will change? | After the analysis, while people still remember | A written record with owned action items |
That three-way split is my own framing, and it matches how Google describes the document. Google’s SRE book on postmortem culture defines a postmortem as “a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause(s), and the follow-up actions to prevent the incident from recurring.” Notice the plural it allows. A failure can need two conditions at once, and the chapter never asks you to pick one.
If you came here for the RCA audio plug or the electronics brand, that is a different RCA; this page is about software.
RCA in testing: root cause analysis in software testing
RCA in software testing is run on a defect that escaped. It asks two questions: where was the defect introduced, and why did the checks not catch it before a user did. My working rule: every escaped defect ends with one new test or check that would have caught it.
In a QA team, root cause analysis in testing starts from a bug report rather than an outage, and the first question sorts the defect by where it came in. The table below is my own framing, not a standard.
| Where the defect came in | The question to ask | The fix that prevents the class |
|---|---|---|
| Requirements | Did anyone say what should happen in this case? | Write the expected behavior into the ticket, with the edge case named |
| Design | Can the flow handle this case at all? | Change the flow, then add a test for the case |
| Code | Is this a plain mistake in the logic? | Fix it and add a test that fails without the fix |
| Environment or configuration | Does it work locally and fail with production settings? | Make staging use production-like settings and check config at deploy |
| Test gap | Does any test cover this path? | Add the missing test to the suite that runs before deploy |
The second question matters more to me than the first. A defect in the code is one bug; a check that could not have caught it will let the next one through as well.
For AI-built apps, the test gap is the answer I would check first. Only 1 of the 21 third-party apps I audited, a healthcare FHIR hub, was credited with a real test suite, and even it skipped sign-up, login and the payment webhook. That 1 of 21 figure comes from the same June and July 2026 audits, a group of apps I chose rather than sampled, so read it as what those apps showed and not as a rate. Flaky test triage, where a test fails at random with no code change, is a different job and not RCA.
Why it matters for a small SaaS: the same outage twice
The cost of skipping RCA shows up the second time. Here are the three patterns I would expect after an outage that nobody analyzed, with what each one leads to. This is my reading, not a measured result.
| What the team did after the outage | What happened next |
|---|---|
| Restarted it and moved on | The same failure returns, because nothing that allowed it has changed |
| Fixed the one line of code | The same class of bug stays in sibling code nobody looked at |
| Blamed the last person to touch it, or the AI tool | People stop reporting near misses, and the conditions stay |
A founder’s app goes down, a restart brings it back, and the team moves on without writing anything down. Weeks later the same failure returns at a busier hour, and with no timeline, no record of the first time and no action item, the investigation starts from nothing. Google’s SRE chapter puts the risk plainly: without “some formalized process of learning from these incidents in place, they may recur ad infinitum.” The lesson I take from it is that a restart ends the outage, and only a written review with owned actions ends the conditions behind it.
As I read it, small teams skip RCA for two reasons: no time, and no data. The data gap is the missing error tracking at the top of this page, and deploys have the same gap. At least 17 of the 21 third-party apps had no deploy gate, and so did all 5 of my own apps: every push ships straight to production with nothing checking it first. The 21 apps and my own 5 belong to the same June and July 2026 audits, a chosen set of apps rather than a sample, so they say nothing about the share of all apps. With no tracker, the errors an RCA would start from are never recorded, and with no gate, nothing checked the change before it shipped.
When a customer or an investor asks for the RCA, what they want, in my reading, is a short written account: what broke, why it was possible, and what changed. That is the summary at the top of a postmortem, and the template below starts with it.
This page starts once the app is back up. For the outage itself, read what to do first when your app is down. If the incident involved customer data or an attacker, the response has legal and notification steps that belong in an incident response plan template and its incident record.
How it works: the methods, the postmortem and the after-action report
This part has four pieces: the methods borrowed from quality engineering, the document, the two rules that keep its actions alive, and the other name that document goes by.
Root cause analysis engineering borrowed from quality work: five whys and the fishbone
Root cause analysis in quality work offers two methods that suit a small software team. Five whys asks why until the answer is a condition you can change. A fishbone diagram sorts possible causes into branches first, so the team does not chase a single chain.
The clearest descriptions I know are two short tools from the US Centers for Medicare & Medicaid Services, written for healthcare quality work rather than software. CMS’s five whys tool describes “drilling down” by asking “Why?” or “What caused this problem?”, and gives a test for when to stop: “If the most recent response were corrected, is it likely the problem would recur?” If the answer is yes, the tool says it is likely a contributing factor, not a root cause, and the team keeps asking.
CMS’s fishbone diagram tool puts the problem at the head of the fish and lists possible causes on the bones, grouped by category, to “keep the team focused on the causes of the problem, rather than the symptoms.” Its major categories often include equipment or supply, environment, rules, policy and procedures, and people, which suit a care setting. For software I’d use code, configuration, data, dependencies, process and monitoring, which is my adaptation, not theirs.
| Method | How it works | Use it when | The trap |
|---|---|---|---|
| Five whys | Ask why the problem happened, then why that answer happened, until fixing the answer would stop a repeat | One failure with a fairly clear chain of events | Stopping at a person, or at the trigger |
| Fishbone diagram | Write the problem at the head, list possible causes under each branch, then ask why about each one | Several things went wrong at once, or the team disagrees on the cause | Filling every branch for the sake of it instead of checking which causes are real |
Here is a worked five whys on a payment failure. It is an illustration of a failure class that AI-built apps can have, not a client’s incident.
- Problem: customers were charged and their accounts were not upgraded.
- Why? The payment webhook handler returned an error.
- Why? It ran out of time waiting on a slow third-party API it called before saving the upgrade.
- Why? That call had no timeout and no retry of its own.
- Why? Nothing in code review or the tests exercises a slow dependency.
- Why? There is no staging drill for a dependency that fails.
The actions come from the last three answers: a timeout and a retry on the call, a test that simulates a slow response, and a staging drill. Making the handler stop erroring this once fixes the first answer and leaves every condition in place.
Three traps are common, in my reading. The first is stopping at a person (the developer forgot a timeout), which names no condition anyone can change. The second is following a single chain when two conditions had to be true together; the fishbone tool itself notes that “there can be more than one root cause.” The third is treating five as a magic number, when the five whys tool says “It often takes three to five whys, but it can take more than five!” Fault trees, Pareto charts and FMEA are other methods you will see named; FMEA comes up again in the questions at the end.
Post mortem software teams write: the nine-part template
A software postmortem is the written record of a root cause analysis, in 9 parts: summary, impact and severity, a UTC timeline, detection, trigger and root causes, what went well, what went badly, action items, and open questions. My working rule is one after every customer-facing outage, any data loss and any near miss.
Copy this into a shared document and fill it in top to bottom.
| Part | What goes in it | The mistake to avoid |
|---|---|---|
| Summary | Three sentences a customer could read: what broke, who it hit, what changed | Writing it for engineers only |
| Impact and severity | Who was affected, for how long, what it cost, and the severity level | Guessing the duration instead of reading it from the logs |
| Timeline, in UTC | Every event with its time, built from logs, deploy history and chat | Writing it from memory |
| Detection | How you found out, and how long after the start | Leaving out that a customer told you first |
| Trigger and root causes | The trigger, then the root causes from the five whys | Listing the trigger as the root cause |
| What went well | What shortened the outage, and any luck that helped | Skipping it, when it is the list of things to keep |
| What went badly or slowly | What made it longer or harder to fix | Turning it into a list of names |
| Action items | One per finding, with an owner, a due date, the change and its verification | Actions with no owner or no date |
| Open questions | What you still do not know, and who will find out | Pretending the story is complete |
The impact part needs a severity scale agreed before the outage, so settle what Sev1 means and the levels below it once and reuse the level here. The timeline depends on your logs carrying times you can trust, which is where error logging best practices earn their keep.
The structure follows the SRE book’s example postmortem in outline. Its headings include Summary, Impact, Root Causes, Trigger, Resolution, Detection, Action Items, Lessons Learned (split into “What went well”, “What went wrong” and “Where we got lucky”) and a Timeline with “all times UTC”. Its root causes note suggests the same method as above: “It’s often helpful to use a technique such as the 5 Whys [Ohn88] to understand the contributing factors.” Mine drops Google’s separate Resolution and Supporting information sections, puts severity inside impact, gives open questions their own part, and gives each action a due date and a way to verify it, where Google’s action table has columns for type, owner and bug.
On when to write one, Google’s chapter lists triggers such as “Data loss of any kind” and “A monitoring failure (which usually implies manual incident discovery)”, and says to define the criteria before an incident happens. I’d add timing: write it within a few working days, while people still remember the order of events.
Post-mortem software is also sold as a product: incident tools with postmortem features built in. In my view, a shared document does the job for a team of five; Google’s own postmortems are “Google Docs, with an in-house template”. A project post-mortem, looking back at a whole project, is a different document from this one.
On a team with no site reliability engineer, the first postmortem does not wait for a hire: the person who led the fix writes the draft, and one other person reviews it. What the full role covers, and what a small team does instead, is in what is a site reliability engineer. The review matters, because in the SRE chapter’s words, “An unreviewed postmortem might as well never have existed.”
Two rules that make the fix stick: blameless, and one action per finding
A postmortem makes a fix stick when it follows two rules. It is blameless: it names systems and roles, never people, so reporting continues. And, as my working rule, every finding gets exactly one action with an owner, a due date, the change to make, and the way it will be verified.
Rule 1 is blameless, in writing. The document names the deploy script and the on-call role, never a person, and for each mistake it asks what about the system made that mistake easy. The SRE chapter gives the reason: “If a culture of finger pointing and shaming individuals or teams for doing the ‘wrong’ thing prevails, people will not bring issues to light for fear of punishment.” Blaming the AI tool fails the same test. An assistant that wrote a bad migration is the trigger; the root cause is that nothing checked the migration before it reached production.
Rule 2 is my working rule: every finding gets exactly one action, and every action has four fields. The findings name nobody, and each action names exactly one person, which is how blameless and accountable fit together.
| Field | What it holds | Example, from the webhook above |
|---|---|---|
| Owner | One named person, not a team | The engineer who looks after billing |
| Due date | A calendar date, not a promise to get to it | The date of the next planned release |
| Change | The specific change to make | Add a timeout and a retry to the third-party call in the webhook handler |
| Verification | The test that now fails without the fix, or the alert that now fires | A test that simulates a slow response passes, and fails when the timeout is removed |
A finding you decide not to act on is written down as an accepted risk, with the name of whoever accepted it. The actions go into the same tracker as product work, because a separate list tends to lose to the roadmap. I’d review open actions at the next incident and about once a month, and I judge a postmortem by how many of its actions got closed, not by how well it reads.
Incident response after action report: the same document under another name
An incident response after-action report is the same review under a name from the US Army, which developed the concept. An AAR answers four questions: what was expected to happen, what actually occurred, what went well and why, and what can be improved and how. A nine-part postmortem answers a request for one with the headings renamed.
The four questions come from USAID’s after-action review guide, which USAID built using the US Army’s TC 25-20 as a guide, and which says an AAR “does not grade success or failure.” The mapping is short: expected and actual are your summary and timeline, what went well has its own part, and what can be improved becomes what went badly plus the action items.
In security incident response, NIST SP 800-61 Rev. 3 (April 2025) maps the old Post-Incident Activity phase to the Improvement category (ID.IM) of its Identify function. Its note on ID.IM-03 says improvements are often identified “when creating follow-up reports for incidents or holding ‘lessons learned’ meetings when an incident’s recovery efforts are concluding, especially if the incident was major.”
When a customer’s contract or a security framework asks for an AAR, the postmortem you already wrote is the document to send, with the four questions as its headings. A security incident adds legal and notification steps that are outside this page and belong to the incident response plan mentioned above.
How to check your own app
A closed incident passes 5 checks: the timeline can be rebuilt from logs alone, every action has an owner, date and verification, the trigger re-created in staging no longer causes the failure, a test error reaches the tracker and raises an alert, and an outsider can repeat the summary.
Run them on your last real incident, or on a staged one if you have not had one yet. Each check can fail, and each leaves evidence worth keeping. The checks are mine, not a standard.
- 01 Rebuild the timeline from logs and deploy history alone. If you cannot place the start time to within a few minutes, the first action is better logging or longer log retention on your host. Keep the timeline with its sources.
- 02 Open the postmortem's action list. Every action has an owner, a due date and a verification, and each closed one links to the change that closed it. Keep the list.
- 03 Re-create the trigger in staging: stop the dependency, push the bad config, or send the event again with the provider's own resend tool rather than re-posting a saved signed request. Confirm the fix holds. Keep the staging run and its date.
- 04 In staging, raise one test error and confirm it reaches the error tracker with the right environment, and that an alert arrives on whatever channel your tracker's plan offers. Keep the event and the alert.
- 05 Hand the summary to someone outside the team and ask them what broke and what changed. Keep their answer.
For a Stripe webhook, the resend in check 3 is the Resend button on the event in the Stripe Dashboard, which works for up to 15 days after the event was created, or the stripe events resend command in the Stripe CLI, which works for up to 30 days.
In the Production Hardening Sprint, two deliverables have written verification lines of their own. Deliverable 6.10, staging outage simulation, is verified this way: “Record outage scenarios, queue behavior, customer messages, and recovery results.” Deliverable 6.3, error tracking, is verified this way: “Send a test error and verify symbolication, environment attribution, and alert delivery.”
Where the sprint fits
In the sprint, deliverable 6.3, error tracking, installs error tracking such as Sentry with protected source maps, environment labels, and alerts. Deliverable 6.10, staging outage simulation, simulates AI or payment-provider failure in staging and verifies the expected recovery behavior. Deliverable 13.3, operating runbooks, documents deployment, rollback, key rotation, backup restoration, and the response to each operational alert. Post-handover support is 14 calendar days of fixes for defects in the delivered sprint work. Every deliverable and its verification is listed in the published scope.
Common questions about root cause analysis
What is the main difference between FMEA and RCA?
FMEA looks forward at how a process could fail, and RCA looks back at why something did fail. CMS’s FMEA guidance describes FMEA as a way to address potential failures “before an adverse event occurs”, and RCA as “a structured way to address problems after they occur.”
In a small SaaS, an FMEA is something you do before launching a flow such as checkout: walk each step and ask how it could break. An RCA is what you do after one did.
What are the 5 steps of root cause analysis?
One workable set of five steps is: define the problem, collect the data, find the causes, make the changes, and verify that the changes hold. Step counts vary from author to author, and the international standard IEC 62740:2015 specifies the steps an RCA process should include. The one I would never drop is the last, because a change nobody verified is still a guess.
What is RCA in devops?
RCA in devops is the same review run after an incident or a failed deploy, written up as a blameless postmortem, with its actions landing in the same backlog as feature work. Google’s SRE book calls blameless postmortems “a tenet of SRE culture.”
What are the common RCA mistakes?
The common RCA mistakes are stopping at a person, stopping at the trigger, following one chain when two conditions were needed, writing actions with no owner, and never re-creating the trigger to prove the fix. Each one leaves the conditions behind the outage in place while the document looks finished.
What is post mortem in cyber security?
A post mortem in cyber security is the review after a security incident, which NIST SP 800-61 Rev. 3 treats as lessons learned that feed its Improvement category. It is the same document as an outage postmortem, plus evidence handling and notification duties that belong to the incident response plan.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase