Seventeen third-party apps out of the 21 I audited had no error tracking or alerting: when a user hits an error, nothing records it. The RCA meaning software teams use is root cause analysis: working back from what users saw to the conditions that allowed it, then changing those conditions. The postmortem is where it gets written down.

The RCA meaning software teams use: root cause analysis, in plain words

RCA in software means root cause analysis: the work done after service is restored to find why a failure was possible, not only what broke. It separates the symptom users saw, the trigger that set it off, the root causes that allowed it, and the contributing factors that slowed the response.

Those 21 third-party apps are the ones I audited in June and July 2026, a selected set of audited apps, not a random sample and not a rate for AI-built apps in general. Without an error tracker, the first record of a failure may be a customer’s email, and an investigation that starts from an email starts late.

RCA stands for root cause analysis, and in software development it begins once customers can use the app again. It is one part of hardening SaaS applications for resilience: the part that stops the same failure from coming back.

In my framing, an RCA keeps four words apart. The symptom is what users saw, such as a checkout that spins forever or an error page. The trigger is what set it off: a deploy, a traffic spike, an expired key. The root causes are the conditions that let the trigger do damage, such as no timeout on an outside call, no test on the path, no alert on the error. Contributing factors are whatever made it slower to notice or fix. A postmortem that lists only the trigger has described the outage without explaining it.

ActivityThe question it answersWhen it happensWhat comes out of it
DebuggingWhat is broken right now?During the outage, or while the bug is openA fix, a rollback or a workaround
Root cause analysisWhy was this possible at all?After service is restoredThe root causes and the conditions to change
PostmortemWhat happened, why, and what will change?After the analysis, while people still rememberA written record with owned action items

That three-way split is my own framing, and it matches how Google describes the document. Google’s SRE book on postmortem culture defines a postmortem as “a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause(s), and the follow-up actions to prevent the incident from recurring.” Notice the plural it allows. A failure can need two conditions at once, and the chapter never asks you to pick one.

If you came here for the RCA audio plug or the electronics brand, that is a different RCA; this page is about software.

RCA in testing: root cause analysis in software testing

RCA in software testing is run on a defect that escaped. It asks two questions: where was the defect introduced, and why did the checks not catch it before a user did. My working rule: every escaped defect ends with one new test or check that would have caught it.

In a QA team, root cause analysis in testing starts from a bug report rather than an outage, and the first question sorts the defect by where it came in. The table below is my own framing, not a standard.

Where the defect came inThe question to askThe fix that prevents the class
RequirementsDid anyone say what should happen in this case?Write the expected behavior into the ticket, with the edge case named
DesignCan the flow handle this case at all?Change the flow, then add a test for the case
CodeIs this a plain mistake in the logic?Fix it and add a test that fails without the fix
Environment or configurationDoes it work locally and fail with production settings?Make staging use production-like settings and check config at deploy
Test gapDoes any test cover this path?Add the missing test to the suite that runs before deploy

The second question matters more to me than the first. A defect in the code is one bug; a check that could not have caught it will let the next one through as well.

For AI-built apps, the test gap is the answer I would check first. Only 1 of the 21 third-party apps I audited, a healthcare FHIR hub, was credited with a real test suite, and even it skipped sign-up, login and the payment webhook. That 1 of 21 figure comes from the same June and July 2026 audits, a group of apps I chose rather than sampled, so read it as what those apps showed and not as a rate. Flaky test triage, where a test fails at random with no code change, is a different job and not RCA.

Why it matters for a small SaaS: the same outage twice

The cost of skipping RCA shows up the second time. Here are the three patterns I would expect after an outage that nobody analyzed, with what each one leads to. This is my reading, not a measured result.

What the team did after the outageWhat happened next
Restarted it and moved onThe same failure returns, because nothing that allowed it has changed
Fixed the one line of codeThe same class of bug stays in sibling code nobody looked at
Blamed the last person to touch it, or the AI toolPeople stop reporting near misses, and the conditions stay

A founder’s app goes down, a restart brings it back, and the team moves on without writing anything down. Weeks later the same failure returns at a busier hour, and with no timeline, no record of the first time and no action item, the investigation starts from nothing. Google’s SRE chapter puts the risk plainly: without “some formalized process of learning from these incidents in place, they may recur ad infinitum.” The lesson I take from it is that a restart ends the outage, and only a written review with owned actions ends the conditions behind it.

As I read it, small teams skip RCA for two reasons: no time, and no data. The data gap is the missing error tracking at the top of this page, and deploys have the same gap. At least 17 of the 21 third-party apps had no deploy gate, and so did all 5 of my own apps: every push ships straight to production with nothing checking it first. The 21 apps and my own 5 belong to the same June and July 2026 audits, a chosen set of apps rather than a sample, so they say nothing about the share of all apps. With no tracker, the errors an RCA would start from are never recorded, and with no gate, nothing checked the change before it shipped.

When a customer or an investor asks for the RCA, what they want, in my reading, is a short written account: what broke, why it was possible, and what changed. That is the summary at the top of a postmortem, and the template below starts with it.

This page starts once the app is back up. For the outage itself, read what to do first when your app is down. If the incident involved customer data or an attacker, the response has legal and notification steps that belong in an incident response plan template and its incident record.

How it works: the methods, the postmortem and the after-action report

This part has four pieces: the methods borrowed from quality engineering, the document, the two rules that keep its actions alive, and the other name that document goes by.

Root cause analysis engineering borrowed from quality work: five whys and the fishbone

Root cause analysis in quality work offers two methods that suit a small software team. Five whys asks why until the answer is a condition you can change. A fishbone diagram sorts possible causes into branches first, so the team does not chase a single chain.

The clearest descriptions I know are two short tools from the US Centers for Medicare & Medicaid Services, written for healthcare quality work rather than software. CMS’s five whys tool describes “drilling down” by asking “Why?” or “What caused this problem?”, and gives a test for when to stop: “If the most recent response were corrected, is it likely the problem would recur?” If the answer is yes, the tool says it is likely a contributing factor, not a root cause, and the team keeps asking.

CMS’s fishbone diagram tool puts the problem at the head of the fish and lists possible causes on the bones, grouped by category, to “keep the team focused on the causes of the problem, rather than the symptoms.” Its major categories often include equipment or supply, environment, rules, policy and procedures, and people, which suit a care setting. For software I’d use code, configuration, data, dependencies, process and monitoring, which is my adaptation, not theirs.

MethodHow it worksUse it whenThe trap
Five whysAsk why the problem happened, then why that answer happened, until fixing the answer would stop a repeatOne failure with a fairly clear chain of eventsStopping at a person, or at the trigger
Fishbone diagramWrite the problem at the head, list possible causes under each branch, then ask why about each oneSeveral things went wrong at once, or the team disagrees on the causeFilling every branch for the sake of it instead of checking which causes are real

Here is a worked five whys on a payment failure. It is an illustration of a failure class that AI-built apps can have, not a client’s incident.

  • Problem: customers were charged and their accounts were not upgraded.
  • Why? The payment webhook handler returned an error.
  • Why? It ran out of time waiting on a slow third-party API it called before saving the upgrade.
  • Why? That call had no timeout and no retry of its own.
  • Why? Nothing in code review or the tests exercises a slow dependency.
  • Why? There is no staging drill for a dependency that fails.

The actions come from the last three answers: a timeout and a retry on the call, a test that simulates a slow response, and a staging drill. Making the handler stop erroring this once fixes the first answer and leaves every condition in place.

Three traps are common, in my reading. The first is stopping at a person (the developer forgot a timeout), which names no condition anyone can change. The second is following a single chain when two conditions had to be true together; the fishbone tool itself notes that “there can be more than one root cause.” The third is treating five as a magic number, when the five whys tool says “It often takes three to five whys, but it can take more than five!” Fault trees, Pareto charts and FMEA are other methods you will see named; FMEA comes up again in the questions at the end.

Post mortem software teams write: the nine-part template

A software postmortem is the written record of a root cause analysis, in 9 parts: summary, impact and severity, a UTC timeline, detection, trigger and root causes, what went well, what went badly, action items, and open questions. My working rule is one after every customer-facing outage, any data loss and any near miss.

Copy this into a shared document and fill it in top to bottom.

PartWhat goes in itThe mistake to avoid
SummaryThree sentences a customer could read: what broke, who it hit, what changedWriting it for engineers only
Impact and severityWho was affected, for how long, what it cost, and the severity levelGuessing the duration instead of reading it from the logs
Timeline, in UTCEvery event with its time, built from logs, deploy history and chatWriting it from memory
DetectionHow you found out, and how long after the startLeaving out that a customer told you first
Trigger and root causesThe trigger, then the root causes from the five whysListing the trigger as the root cause
What went wellWhat shortened the outage, and any luck that helpedSkipping it, when it is the list of things to keep
What went badly or slowlyWhat made it longer or harder to fixTurning it into a list of names
Action itemsOne per finding, with an owner, a due date, the change and its verificationActions with no owner or no date
Open questionsWhat you still do not know, and who will find outPretending the story is complete

The impact part needs a severity scale agreed before the outage, so settle what Sev1 means and the levels below it once and reuse the level here. The timeline depends on your logs carrying times you can trust, which is where error logging best practices earn their keep.

The structure follows the SRE book’s example postmortem in outline. Its headings include Summary, Impact, Root Causes, Trigger, Resolution, Detection, Action Items, Lessons Learned (split into “What went well”, “What went wrong” and “Where we got lucky”) and a Timeline with “all times UTC”. Its root causes note suggests the same method as above: “It’s often helpful to use a technique such as the 5 Whys [Ohn88] to understand the contributing factors.” Mine drops Google’s separate Resolution and Supporting information sections, puts severity inside impact, gives open questions their own part, and gives each action a due date and a way to verify it, where Google’s action table has columns for type, owner and bug.

On when to write one, Google’s chapter lists triggers such as “Data loss of any kind” and “A monitoring failure (which usually implies manual incident discovery)”, and says to define the criteria before an incident happens. I’d add timing: write it within a few working days, while people still remember the order of events.

Post-mortem software is also sold as a product: incident tools with postmortem features built in. In my view, a shared document does the job for a team of five; Google’s own postmortems are “Google Docs, with an in-house template”. A project post-mortem, looking back at a whole project, is a different document from this one.

On a team with no site reliability engineer, the first postmortem does not wait for a hire: the person who led the fix writes the draft, and one other person reviews it. What the full role covers, and what a small team does instead, is in what is a site reliability engineer. The review matters, because in the SRE chapter’s words, “An unreviewed postmortem might as well never have existed.”

Two rules that make the fix stick: blameless, and one action per finding

A postmortem makes a fix stick when it follows two rules. It is blameless: it names systems and roles, never people, so reporting continues. And, as my working rule, every finding gets exactly one action with an owner, a due date, the change to make, and the way it will be verified.

Rule 1 is blameless, in writing. The document names the deploy script and the on-call role, never a person, and for each mistake it asks what about the system made that mistake easy. The SRE chapter gives the reason: “If a culture of finger pointing and shaming individuals or teams for doing the ‘wrong’ thing prevails, people will not bring issues to light for fear of punishment.” Blaming the AI tool fails the same test. An assistant that wrote a bad migration is the trigger; the root cause is that nothing checked the migration before it reached production.

Rule 2 is my working rule: every finding gets exactly one action, and every action has four fields. The findings name nobody, and each action names exactly one person, which is how blameless and accountable fit together.

FieldWhat it holdsExample, from the webhook above
OwnerOne named person, not a teamThe engineer who looks after billing
Due dateA calendar date, not a promise to get to itThe date of the next planned release
ChangeThe specific change to makeAdd a timeout and a retry to the third-party call in the webhook handler
VerificationThe test that now fails without the fix, or the alert that now firesA test that simulates a slow response passes, and fails when the timeout is removed

A finding you decide not to act on is written down as an accepted risk, with the name of whoever accepted it. The actions go into the same tracker as product work, because a separate list tends to lose to the roadmap. I’d review open actions at the next incident and about once a month, and I judge a postmortem by how many of its actions got closed, not by how well it reads.

Incident response after action report: the same document under another name

An incident response after-action report is the same review under a name from the US Army, which developed the concept. An AAR answers four questions: what was expected to happen, what actually occurred, what went well and why, and what can be improved and how. A nine-part postmortem answers a request for one with the headings renamed.

The four questions come from USAID’s after-action review guide, which USAID built using the US Army’s TC 25-20 as a guide, and which says an AAR “does not grade success or failure.” The mapping is short: expected and actual are your summary and timeline, what went well has its own part, and what can be improved becomes what went badly plus the action items.

In security incident response, NIST SP 800-61 Rev. 3 (April 2025) maps the old Post-Incident Activity phase to the Improvement category (ID.IM) of its Identify function. Its note on ID.IM-03 says improvements are often identified “when creating follow-up reports for incidents or holding ‘lessons learned’ meetings when an incident’s recovery efforts are concluding, especially if the incident was major.”

When a customer’s contract or a security framework asks for an AAR, the postmortem you already wrote is the document to send, with the four questions as its headings. A security incident adds legal and notification steps that are outside this page and belong to the incident response plan mentioned above.

How to check your own app

A closed incident passes 5 checks: the timeline can be rebuilt from logs alone, every action has an owner, date and verification, the trigger re-created in staging no longer causes the failure, a test error reaches the tracker and raises an alert, and an outsider can repeat the summary.

Run them on your last real incident, or on a staged one if you have not had one yet. Each check can fail, and each leaves evidence worth keeping. The checks are mine, not a standard.

  1. 01 Rebuild the timeline from logs and deploy history alone. If you cannot place the start time to within a few minutes, the first action is better logging or longer log retention on your host. Keep the timeline with its sources.
  2. 02 Open the postmortem's action list. Every action has an owner, a due date and a verification, and each closed one links to the change that closed it. Keep the list.
  3. 03 Re-create the trigger in staging: stop the dependency, push the bad config, or send the event again with the provider's own resend tool rather than re-posting a saved signed request. Confirm the fix holds. Keep the staging run and its date.
  4. 04 In staging, raise one test error and confirm it reaches the error tracker with the right environment, and that an alert arrives on whatever channel your tracker's plan offers. Keep the event and the alert.
  5. 05 Hand the summary to someone outside the team and ask them what broke and what changed. Keep their answer.

For a Stripe webhook, the resend in check 3 is the Resend button on the event in the Stripe Dashboard, which works for up to 15 days after the event was created, or the stripe events resend command in the Stripe CLI, which works for up to 30 days.

In the Production Hardening Sprint, two deliverables have written verification lines of their own. Deliverable 6.10, staging outage simulation, is verified this way: “Record outage scenarios, queue behavior, customer messages, and recovery results.” Deliverable 6.3, error tracking, is verified this way: “Send a test error and verify symbolication, environment attribution, and alert delivery.”

Where the sprint fits

In the sprint, deliverable 6.3, error tracking, installs error tracking such as Sentry with protected source maps, environment labels, and alerts. Deliverable 6.10, staging outage simulation, simulates AI or payment-provider failure in staging and verifies the expected recovery behavior. Deliverable 13.3, operating runbooks, documents deployment, rollback, key rotation, backup restoration, and the response to each operational alert. Post-handover support is 14 calendar days of fixes for defects in the delivered sprint work. Every deliverable and its verification is listed in the published scope.

Common questions about root cause analysis

What is the main difference between FMEA and RCA?

FMEA looks forward at how a process could fail, and RCA looks back at why something did fail. CMS’s FMEA guidance describes FMEA as a way to address potential failures “before an adverse event occurs”, and RCA as “a structured way to address problems after they occur.”

In a small SaaS, an FMEA is something you do before launching a flow such as checkout: walk each step and ask how it could break. An RCA is what you do after one did.

What are the 5 steps of root cause analysis?

One workable set of five steps is: define the problem, collect the data, find the causes, make the changes, and verify that the changes hold. Step counts vary from author to author, and the international standard IEC 62740:2015 specifies the steps an RCA process should include. The one I would never drop is the last, because a change nobody verified is still a guess.

What is RCA in devops?

RCA in devops is the same review run after an incident or a failed deploy, written up as a blameless postmortem, with its actions landing in the same backlog as feature work. Google’s SRE book calls blameless postmortems “a tenet of SRE culture.”

What are the common RCA mistakes?

The common RCA mistakes are stopping at a person, stopping at the trigger, following one chain when two conditions were needed, writing actions with no owner, and never re-creating the trigger to prove the fix. Each one leaves the conditions behind the outage in place while the document looks finished.

What is post mortem in cyber security?

A post mortem in cyber security is the review after a security incident, which NIST SP 800-61 Rev. 3 treats as lessons learned that feed its Improvement category. It is the same document as an outage postmortem, plus evidence handling and notification duties that belong to the incident response plan.