Which 5 numbers tell you in one look whether the business is healthy? Uptime, error rate, latency, signups and revenue. How to build an ops dashboard is mostly a matter of pulling those five out of the tools that already hold them, putting them on one screen, and checking each against its source.
What is an operations dashboard?
An operations dashboard is one private screen that shows the current state of a service and its business side by side, read in seconds and opened daily. For a small SaaS it holds 5 numbers: uptime, error rate, latency, signups and revenue. It informs; it does not interrupt like an alert.
It is the part of logging and monitoring for a small SaaS that you go and look at, as opposed to the part that comes to find you. What makes it operational and not analytical is the question it answers: is anything wrong right now, and did yesterday go normally? A question such as why churn moved this quarter belongs on a different screen, one you study for an hour, not glance at over coffee.
Three things it is not. It is not an alert: alerts interrupt someone when a threshold is crossed, and the dashboard waits to be opened. It is not a public status page, which faces customers and has to stay up when the app is down; this screen is for you. And it is not a BI project with a chart for everything. Choosing among every signal an app can emit is the subject of application monitoring best practices for a small SaaS; the dashboard needs only five of them.
A single screen like this is deliverable 8.8, the unified operations dashboard, in the Production Hardening Sprint’s published scope.
What goes wrong without it, and with a bad one
There are two ways to get this wrong. The first is having no single view at all, so the health of the business is pieced together from five tabs, or not at all. The second is a single view that lies: a screen full of green tiles while customers fail to check out.
The first failure can start one step earlier, with a number that has no source. In my audits, 17 of the 21 third-party apps had no error tracking or alerting: when a user hits an error, nothing records it. I audited those 21 apps in June and July 2026: 11 public vibe-coded apps audited exhaustively across all 12 pillars, and 10 held-out apps audited blind. They are a selected set of audited apps, not a random sample, so the count describes them and is not a rate for all AI-built apps. My reading of it for this page: where there is no error tracker, the error tile has nothing to read, so the dashboard is the last step of the work, not the first.
Checking five dashboards every morning
Checking five dashboards every morning means five logins, five date ranges and no way to see that errors rose in the same hour signups fell. The ritual also gets skipped on busy days, which are the days it matters. One screen that reads all five replaces it.
| Tool you open | The number you want from it | What the morning tour misses |
|---|---|---|
| Uptime monitor | Did every check pass since yesterday? | Whether the checked route is the one customers use |
| Error tracker | Did errors go up? | The total request count, so a spike has no denominator |
| Host metrics | Is the app slow? | Which routes are slow, and since when |
| Product analytics | How many people signed up? | That signups dipped in the same hour errors rose |
| Payment provider | Did money come in? | Last Tuesday’s figure, so a quiet day looks normal |
Each tab has its own login, its own default time zone and its own idea of when today starts, so the five answers may not cover the same window. In my reading, the tour misses three things. It misses relationships: errors rising in the hour signups fell only shows up when the two numbers sit next to each other. It misses slow drift, because nobody compares today with last Tuesday from memory. And it misses the days it does not happen at all. The cost is attention, not money, and the fix is not a sixth tool.
The watermelon dashboard: green tiles, unhappy customers
The watermelon effect is a report that is green on the outside and red on the inside: every target met while the people it serves are unhappy. On a SaaS dashboard it looks like a perfect uptime score while checkout fails. Pair each health tile with a customer-outcome number and show when each was last updated.
The term comes from IT service management, the ITIL world of service desks, where a desk can hit every SLA on its report and still leave its users unhappy. When the green number is a contract target, people call it the watermelon SLA effect; it has nothing to do with the fruit.
On a SaaS health screen, watermelon reporting looks like the rows below. These are my reading of how each tile goes green while the customer has a bad day, not cases from a source.
| Green tile | What it measured | What the customer saw | The paired number that would have shown it |
|---|---|---|---|
| Uptime: every check passed | The home page answered | Checkout failed | Completed checkouts in the same hours |
| Error rate low | Requests that returned an error status, while failures came back as success with an empty body | Blank screens | Successful sign-ins, or completed checkouts |
| Average latency fine | The mean of all requests | The slowest requests timed out | 95th percentile latency and the count of timeouts |
| Revenue flat | A cached figure from yesterday | Renewals failing, unseen | Payments received today, with the tile’s last-updated time |
The first row has a known fix: point your uptime monitoring at a route that can fail. The other rows need the pairing rule, which is my working rule for keeping watermelon metrics off the screen. Every health tile sits next to a customer-outcome number that comes from a different system: uptime beside successful sign-ins, error rate beside completed checkouts. And every tile shows when it was last updated, so a dead data source turns gray instead of staying green.
AWS supplied a public example from its own Service Health Dashboard. The cause of its 28 February 2017 Amazon S3 disruption, and the safeguard that answers it, are material for a runbook template; the part that concerns a dashboard is a different paragraph of AWS’s post-event summary of the 2017 S3 disruption. AWS wrote: “From the beginning of this event until 11:37AM PST, we were unable to update the individual services’ status on the AWS Service Health Dashboard (SHD) because of a dependency the SHD administration console has on Amazon S3.” Instead, it used “the AWS Twitter feed (@AWSCloud) and SHD banner text to communicate status”, and afterwards it said “we have changed the SHD administration console to run across multiple AWS regions”.
The lesson I take from it is my opinion, not AWS’s: a status screen that depends on the thing it reports on goes quiet exactly when you need it. That is why every tile on your screen needs a last-updated time and a source that does not share the failure it is meant to show.
How to build an ops dashboard: five numbers, their sources, one screen
An ops dashboard is built in 5 steps: define the five numbers precisely, find the system that already holds each one, choose a build route, lay the tiles out with a comparison and a last-updated time, and state how often each refreshes. Most of the work is definitions, not charts.
The sections below follow that order: the five numbers and their sources, three ways to build the screen, the business numbers that sit beside the health ones, and a checklist for the finished page. Every table is plain text you can copy into a doc or a sheet as a template; there is nothing to download.
The five numbers and where each comes from
The definitions in this table are my working rules. Hold each tile to its row: if you cannot say which row a number on your screen matches, the tile is not finished.
| Tile | Exact definition | Source system | How it is read | Refresh | What bad looks like |
|---|---|---|---|---|---|
| Uptime | Share of successful checks in the last 24 hours and the last 30 days | Your uptime monitor | Its API or export, where it offers one | As often as the monitor checks | A failed check, or a check that only loads the home page |
| Error rate | Failed requests divided by all requests, in the last hour and the last day | The host’s request logs, or the error tracker’s performance data | A log query or the tracker’s API | About a minute | A jump against the same hour last week |
| Latency | 95th percentile response time on key routes (sign-in, checkout, the main page) | The tracker’s performance data, or the host’s metrics | The tracker’s API or the host’s metrics | About a minute | p95 rising week on week, or timeouts appearing |
| Signups | New accounts today and over the trailing 7 days, in a stated time zone | Your own database | The SQL count below | About a minute | Today below the same weekday last week |
| Revenue | Payments received today and month to date, plus MRR, test mode excluded | Your payment provider | Its API for payments; MRR as the provider computes it, or your own calculation | About a minute for payments | A flat line on a weekday, or an MRR figure with no stated definition |
The checks themselves, and which route each one should load, are a separate job: the uptime monitor setup, not the dashboard. Before you plan the uptime and error tiles, check that your monitor and tracker can be read from code. UptimeRobot’s API, for one, returns each monitor’s status, uptime ratio and average response time, and Sentry’s web API can be used to manage and export data. One catch on the error tile, in my reading: error events alone carry no count of all requests, so an error rate taken from a tracker needs its performance data, which may mean turning tracing on. How long the underlying logs are kept limits how far back the error tile can look, which is the question of how long to keep application logs.
Each tile shows three things: the number, a comparison (the same weekday last week is a sensible default), and a last-updated time. The signups count runs against your own database, with both window bounds passed in the dashboard’s time zone:
-- New signups in one closed window, bounds set in the dashboard's time zone
SELECT count(*) AS signups
FROM users
WHERE created_at >= :window_start
AND created_at < :window_end;
Revenue needs one more decision, because MRR has more than one definition. Stripe’s Billing overview has a Configure button for this, in Stripe’s words: “Click Configure to change how Stripe calculates Monthly Recurring Revenue (MRR), Churn, and Active Subscribers.” Stripe also notes that “Subtracting discounts from MRR is considered a more conservative approach to reporting MRR”. So the revenue tile states which definition it uses. Stripe’s billing analytics documentation lists MRR among the metrics in the Dashboard’s Billing overview and describes a CSV export of the billing metrics. For a live tile, Stripe’s Analytics API returns Stripe-defined metrics such as MRR, but it is in private preview, with access by request. Without that access, my reading is that a tile built in code either computes MRR from subscription data under a written definition or shows payments only.
Combine uptime and revenue in one view: three ways to build it
Uptime and revenue are combined in one view 3 ways: a dashboard tool that reads each product’s data, one admin page in your app that queries the database and calls the other tools server-side with read-only keys, or a scheduled script that writes the numbers to a sheet each hour.
| Route | What it takes | When it fits | The catch |
|---|---|---|---|
| A dashboard tool that reads other tools’ data | Connecting each source as a data source, little code | You want the least code, or already run Grafana | Another login; each source needs a connector; check each tool’s own pricing page |
| One admin page inside the app | A server-side route, read-only keys, a short cache | An app a builder generated, where you already sign in | You maintain it, and the keys must never reach the browser |
| A scheduled script that writes to a sheet | A timed job, the same API calls, a spreadsheet | The first month | Only as fresh as its schedule, and crude; stale sources only show if the script writes a timestamp |
Route one is a tool like Grafana, an open-source option, where a data source is “a connection to a storage backend that holds your data, such as a Prometheus server, a Loki instance, a SQL database, or a cloud monitoring service.” Its built-in list includes PostgreSQL, so if your database is PostgreSQL the signups count can run straight against it, and Grafana’s data sources documentation points to a plugin catalog for sources outside that list. Whether a plugin exists for your monitor and payment provider is the thing to check before choosing this route. Hosted dashboard products sit in the same route; check their connectors the same way.
Route two is one admin page in your own app. A server-side route queries your database for signups, calls the monitor’s, the tracker’s and the payment provider’s APIs with read-only keys, and renders five tiles. My working rule is to cache the result for about a minute, so opening the page ten times does not repeat every API call ten times. The page sits behind the admin role check, and the keys stay on the server. For the payment side, Stripe’s restricted API keys let you set each resource’s permission to Read, Write or None, and “The default value for all permissions is None”, so a dashboard key gets Read on the few resources it needs and nothing else.
Where those keys end up matters, because 6 of the 21 third-party apps shipped a real secret. That count comes from the same June and July 2026 audits as the error-tracking figure above, a chosen group of apps rather than a sample, so it speaks for those 21 only. This route is the best fit for an app made with a builder, in my reading: the builder can generate the page, and your job is to review where the keys live.
Route three is a scheduled script that writes the five numbers to a sheet each hour. It is crude and dependable, and it is good for the first month while you learn which tiles you actually look at. Enterprise BI tools such as Tableau or Power BI are more than five tiles need. My working rule for choosing between the three: pick the route you will still open in three months.
The business numbers on the same screen: a SaaS metrics cheat sheet
A SaaS metrics cheat sheet for an operations screen has 7 business numbers: new signups, activation, trial conversion, MRR, net new MRR, customer churn and failed payments awaiting retry. Each gets one written definition and one source system. CAC, LTV and the rule of 40 I keep off it as monthly board numbers.
These are SaaS operations metrics, the ones on a SaaS metrics dashboard that can move between Monday and Tuesday, not a finance pack. Where Stripe’s documentation defines a metric, the row uses Stripe’s definition; the others are my working definitions.
| Metric | Definition | Source system | How often it changes |
|---|---|---|---|
| New signups | New accounts created in the window (my definition) | Your database | Daily |
| Activation | Share of new signups who complete the first core action of the product (my definition) | Your product analytics or database | Daily |
| Trial conversion | Stripe: trials that converted to a paid plan in the last 30 days, divided by trials that ended in the last 30 days | Payment provider | Daily |
| MRR | Stripe: the sum of the monthly-normalized value of all active and past_due subscriptions, excluding taxes, free plans and metered products | Payment provider | Daily |
| Net new MRR | Stripe calls it MRR growth: new, reactivation and expansion MRR, minus contraction and churn MRR, adjusted for currency | Payment provider | Weekly |
| Customer churn | Stripe’s subscriber churn rate: churned subscribers in the past 30 days, divided by active subscribers 30 days ago plus new subscribers in the past 30 days | Payment provider | Weekly |
| Failed payments awaiting retry | Payments that failed and are still being retried (my definition) | Payment provider | Daily |
Activation needs the first core action to be recorded as an event, one of the four product events worth instrumenting. The last row is the early sign of failed payments and involuntary churn, and in my reading it is the business number easiest to miss without a tile of its own.
CAC, LTV, payback period and the rule of 40 I treat as board numbers: reviewed monthly, slow to move, and noise on a daily screen. If you went looking for best-in-class SaaS metrics to hold yours against, my reading is that benchmarks shift with a company’s stage and with whoever collected them, so this page prints none. One more check belongs here: the revenue tile has to agree with your app’s own subscription rows, and the job that does it is keeping Stripe subscriptions in sync with your database.
A founder metrics dashboard checklist
A founder metrics dashboard checklist has 9 lines. The ones that keep the screen honest: every tile has a written definition and a named source, every tile shows when it was last updated and goes gray when stale, and the numbers were checked against source this month.
- Five health and business tiles first, and no more than about ten in total.
- Every tile has a written definition.
- Every tile names its source system.
- Every tile shows a comparison, such as the same weekday last week.
- Every tile shows its last-updated time and goes gray when stale.
- Each health tile sits beside a customer-outcome number from a different system.
- One time zone, stated on the screen.
- The page is behind admin sign-in and uses read-only keys kept on the server.
- Someone opens it daily, and the numbers were checked against source this month.
The design test for the whole page is whether it reads in a few seconds: a big number, a comparison beside it, a gray tile where a source has gone quiet, and nothing that needs a legend.
How to verify it: verify the dashboard numbers match the source
Dashboard numbers are verified against their source over one closed window, such as yesterday in the dashboard’s time zone. Count signups in the database, read revenue in the payment provider, and compare each tile. Then time a test signup against the stated refresh, and break one source to confirm the tile goes gray.
Run the seven checks below to verify the dashboard numbers match the source, once a month and again after any change of tool. They are my own checks, and each one can fail.
- 01 Pick one closed window: yesterday, midnight to midnight, in the time zone the dashboard states.
- 02 Signups: run the count on your users table for that window and compare it with the tile.
- 03 Revenue: open the payment provider's dashboard for the same window and compare gross payments. Write down whether the tile shows gross, net of fees or net of refunds, which MRR definition the provider is configured to use, and confirm test-mode payments are excluded.
- 04 Uptime: compare the tile with the monitor's own report for the same window.
- 05 Error rate and latency: compare each tile with the tracker's or the host's chart for the same window.
- 06 Refresh: create a signup with an address the tile does not filter out, and time how long the tile takes to move against the refresh it claims. Delete the account afterwards.
- 07 Staleness: on staging, revoke or break one source's key and confirm the tile goes gray with its last-updated time, rather than staying green.
If you changed the MRR configuration recently, allow for Stripe’s note that “Changes take 24-48 hours to appear in your configuration” before you compare. Record each comparison in a table like this one:
| Tile | Dashboard value | Source value | Same window? | Difference | Likely cause to check first |
|---|---|---|---|---|---|
| Signups | From the tile | From the SQL count | Yes or no | Tile minus source | Time zone, test accounts, or a different definition of a signup |
| Revenue | From the tile | From the provider’s dashboard | Yes or no | Tile minus source | Test data, refunds, fees or currency |
| Uptime | From the tile | From the monitor’s report | Yes or no | Tile minus source | A different window or check set |
| Error rate | From the tile | From the tracker or host chart | Yes or no | Tile minus source | Cached values |
| Latency | From the tile | From the tracker or host chart | Yes or no | Tile minus source | Average shown where p95 was meant |
The evidence to keep is the filled table with its date, plus the timed refresh from check six.
In the Production Hardening Sprint, deliverable 8.8 is verified this way: “Compare the displayed metrics with their source systems and verify refresh behavior.”
Where the sprint does this
The Production Hardening Sprint’s deliverable 8.8 is to “Create one view of uptime, error rate, latency, signups, and revenue for the application’s relevant services”; its verify line is quoted in the verify section above. The result goes into the production readiness report, which accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Hosting, paid tools, and API usage remain in your accounts. Deliverable 8.8 sits with the rest of the monitoring deliverables in the published scope.
Common questions about operations dashboards
What are the four types of dashboards?
The usual answer is operational, analytical, strategic and tactical, a common convention in business-intelligence writing with no single source behind it. Operational dashboards show what is happening now, analytical ones help explain why something changed, strategic ones track long-range goals, and tactical ones follow a team’s progress over weeks. The screen on this page is operational.
Amazon’s article on building dashboards for operational visibility sorts them another way, by who reads them and how much of the system they cover, with names such as customer experience dashboards and system-level dashboards.
What is the 5 second rule for dashboards?
The 5 second rule is a design heuristic: someone opening a dashboard should get its main message in about five seconds. It is a rule of thumb with no single author. For a five-tile screen it means big numbers, one comparison beside each, and gray for anything stale, so a glance tells you whether to look closer.
Can you give me an example of an operational dashboard?
Yes: five tiles on one admin page, each comparing yesterday with the same weekday last week. Uptime sits beside successful sign-ins, error rate beside completed checkouts, p95 latency on checkout beside the count of timeouts, signups beside activation, and payments received beside failed payments awaiting retry. Every tile carries its last-updated time in one stated time zone, and a tile that has not refreshed turns gray.
What are the 5 most important metrics for SaaS companies?
For a daily operations screen, my pick is signups, activation, MRR, churn and failed payments awaiting retry. I leave the board numbers off that list: CAC, LTV, payback and the rule of 40 move monthly and answer a different question.
What is the rule of 40 for a SaaS company?
The rule of 40 says a SaaS company’s growth rate plus its profit should add up to 40 percent. Brad Feld wrote it up on 3 February 2015 after hearing it at a board meeting from a late-stage investor, in these words: “The 40% rule is that your growth rate + your profit should add up to 40%.” He added that “These are for SaaS companies at scale”, which he put at no less than 50 million dollars in revenue. That is why it is a board number and not a tile on a small SaaS screen.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase