A small SaaS needs 4 signals and one alert channel before it needs an observability platform: request rate, error rate, latency, and whether the app answers from outside. Application monitoring best practices at this size mean wiring those four first, alerting only on what needs a person, and proving each alert fires. Tracing comes later.

Application monitoring best practices: the short list for a small SaaS

Application monitoring best practices for a small SaaS come down to 8 habits: watch user-facing signals first, check from outside too, use one alert channel, alert only on what needs a person, put the first step in every alert, start thresholds loose, test alerts by breaking things, and review the list when the stack changes.

This page sits under the area’s pillar on logging and monitoring, and it covers the watching half for an app with paying users and one or two people running it.

The table puts each habit next to what it means on a managed host and the mistake it prevents. Application performance monitoring best practices for a larger team would add distributed tracing and profiling on top; this page gets to tracing later, because a two-person team needs the eight rows first.

PracticeWhat it means for a small SaaS on a managed hostThe mistake it prevents
Watch user-facing signals before machine signalsTrack requests, errors and response time per route before CPU or memory, which the host already watchesA calm CPU chart while checkout fails
Check from outside as well as insideAn uptime monitor calls the public URL from another networkAn app that reports itself healthy while nobody can reach it
One alert channel a named person readsEvery alert lands in one Slack channel or inbox, and one person owns it this weekAlerts spread across inboxes nobody opens
Alert on what needs a person, log the restAn alert means someone must act now; everything else goes to logs and a daily lookAlert fatigue, then muted alerts
Every alert names its first stepThe message says what to open first: a dashboard, a log query, the provider’s status pageSomeone woken at night guessing where to start
Thresholds start loose and tighten with your own trafficBegin with the starting points further down and adjust after a few weeks of real trafficConstant false alarms in week one
Test every alert by breaking the thingThrow a test error, stop the health route, cross a threshold on purposeAn alert that has never fired and never will
Review the list when the stack changesA new host, queue, payment provider or model call adds a lineMonitoring the app you had last year

None of the eight depends on the framework, so they hold as best practices for monitoring any web app on a managed host. Infrastructure monitoring best practices for a SaaS at this size fit in one short list, database connections, storage, function errors and plan limits, because the provider runs the machines. Buying a production monitoring platform does not change these best practices: at this size they are habits, and no platform supplies them. For a checklist you can run, covering the application performance items and the infrastructure monitoring items above, use the six checks under “How to check your own app” below.

What application performance monitoring means, and what APM traces add

Application performance monitoring measures how long requests take and where the time goes, from inside the app. An APM trace follows one request through the functions, queries and outbound calls it touches as timed spans, which is how a slow page gets pinned to the one query that made it slow.

APM stands for application performance monitoring. Application monitoring itself means knowing whether the app works for its users right now; the performance part adds how long each request took and which step used the time. In OpenTelemetry’s terms, traces “give us the big picture of what happens when a request is made to an application”, and “A span represents a unit of work or operation”: one function call, one database query, one call to another service.

A metric can tell you the pricing page got slower this week. A trace tells you which query made it slow for this user, on this request, and how long that query waited. Web application performance monitoring also covers the browser side, such as page load and Core Web Vitals, which belong to front-end performance work rather than to this page.

Why it matters for a small SaaS: monitoring SaaS applications nobody is watching

In my audits, 17 of the 21 third-party apps had no error tracking or alerting: when a user hits an error, nothing records it.

Those 21 are the public and held-out third-party apps I audited in June and July 2026. They were chosen, not drawn at random, so the count describes that set and is not a rate for AI-built apps in general.

In an AI coding workspace I audited, the code had no error tracking, though the report notes the owner might run a log drain the code does not show. An observability package was a declared dependency imported nowhere, and every API route’s error handling was a bare console line in the host’s short-lived function logs.

My take: a monitoring package in the dependency list is not monitoring. The test is whether an error thrown in production reaches a person.

How to monitor SaaS applications starts with the failures that make no noise. These three are generic shapes, not events from one app:

What brokeWhat the user seesWho notices first without monitoring
The silent one: a webhook or scheduled job stops runningA paid plan that never activates, a reminder that never arrivesA customer, in an email to support
The slow one: one query degrades as a table growsPages that take a little longer every weekA customer who complains, or one who leaves without writing
The partial one: login works, checkout does notA button that spins, or an error on one page onlyA customer, in an email asking why they cannot pay

In IT usage, SaaS monitoring means something else: watching the SaaS tools a company buys, such as its CRM or office suite, which is a different product from the one this page covers.

24/7 managed monitoring for a small SaaS is a service built for teams that have an on-call rota to hand alerts to. With two people, the honest version is an outside check plus one alert channel with stated hours when someone reads it.

The outside check is its own control, with its own tools and test steps, in uptime monitoring for founders.

How it works: the signals, the setup, your stack, your cloud and your model calls

These six parts come in the order a small team usually meets them: which numbers to watch, which pieces collect them, what an agent adds for your language, what your cloud already provides, what to record about model calls, and when tracing is due.

What are RED metrics? RED, USE and the shortest list of signals worth watching

RED metrics are 3 numbers for every service: rate, errors and duration. USE is utilization, saturation and errors per resource. On a managed host the provider already watches the machines, so a small SaaS wires RED first and adds an outside uptime check as the fourth signal.

Tom Wilkie created the RED method in 2015. A 2018 Grafana Labs blog post by Julie Dam quotes his reason: “We really wanted a microservices-oriented monitoring philosophy, so we came up with the RED Method.” The post’s own definition of each of the three is in the table below.

Brendan Gregg’s USE method says: “For every resource, check utilization, saturation, and errors.” A resource there means “all physical server functional components (CPUs, disks, busses, …)”. The four golden signals in Google’s SRE book are “latency, traffic, errors, and saturation.”

MethodWhat it watchesIts termsWhen a small SaaS uses it
RED (Tom Wilkie, quoted in a 2018 Grafana Labs post)Each service that answers requestsRate: “the number of requests per second”; errors: “the number of those requests that are failing”; duration: “the amount of time those requests take”First, per route or endpoint
USE (Brendan Gregg)Each resource: CPUs, disks, bussesUtilization, saturation, errorsMostly the provider’s side on a managed host; your part is the database panels
Four golden signals (Google’s SRE book)A user-facing systemLatency, traffic, errors, saturationWhen a queue or connection pool starts to fill, saturation joins RED

In practice, RED monitoring for a small SaaS means three numbers per route and an alert on the two that hurt users most: errors and duration. Zero traffic is the third alert: a route that normally gets requests and suddenly gets none points at something that broke before the request reached your code.

My working rule for starting points: alert when the error rate stays above about 2 percent for around 5 minutes, when p95 duration stays above about 2 seconds for around 10 minutes, or when requests drop to zero during business hours. These are my starting points, not standards; tighten each one once you have a few weeks of your own traffic to compare against.

The metric types that store these numbers, and which type fits a count or a duration, are covered in Prometheus metric types.

The smallest setup that covers them

The smallest monitoring setup that covers a small SaaS has 4 pieces: an error tracker in the app, the host’s request metrics, an outside uptime check, and one alert channel a named person reads. The first three send their alerts to the channel, and each alert gets tested once by breaking the thing it watches.

Web application performance monitoring tools come in a handful of types: error trackers, APM agents, uptime monitors, log search and metrics dashboards. The table maps each signal to one of those types, never to a product.

SignalWhere it comes from on a managed hostThe piece that collects itStarting alertThe test
ErrorsExceptions thrown in your codeAn error tracker in the appThe error-rate threshold in the RED section aboveCheck 1 below
Rate and durationThe host’s built-in request metrics, if your host shows them on your plan (hosts differ), or an OpenTelemetry exporterThe host dashboard or an OpenTelemetry exporterThe duration and zero-traffic thresholds aboveCheck 3 below
The outside checkRequests to the public URL from another networkAn uptime monitorDown from outside, as the uptime article sets itCheck 2 below
Database connections and slow queriesThe database provider’s metricsThe provider’s own panelsSet from your plan’s connection limitCheck 3 below

How to wire the error tracker is in error logging best practices; this page only needs it to exist and to reach the channel. Every alert goes to one channel; the routing itself is a separate setup, Slack alerting. Logs carry a request id so an alert leads straight to the lines that explain it, which is the core of how to do logging. One screen for all four signals is a separate question: how to build an ops dashboard. Comparing named products belongs to Sentry vs Datadog.

APM for your stack: Node.js, Python, Java and PHP

An APM agent is a library loaded at start-up that times requests, queries and outbound calls without code changes. OpenTelemetry’s zero-code instrumentation covers Java, Python, Node.js and PHP among other languages; for Java it is a Java agent JAR attached to any Java 8+ application. My working rule: add one when “which step is slow” becomes a weekly question.

OpenTelemetry’s zero-code instrumentation pages give each language’s start command and what it captures. They print no status, so the status column below comes from OpenTelemetry’s status page, which rates each language’s SDK by signal.

One thing to add by hand is the same on every stack, by my reading: the request id on every log line and the tenant or account id on spans, and never a secret in either.

StackStatus the docs state (SDK: traces, metrics, logs)How it is startedWhat it captures without code
JavaStable, Stable, StableAdd -javaagent:path/to/opentelemetry-javaagent.jar to the JVM startup argumentsTelemetry at the “edges” of the app: “inbound requests, outbound HTTP calls, database calls, and so on”
Node.jsStable, Stable, Developmentnode --require @opentelemetry/auto-instrumentations-node/register app.js, or the same flag in NODE_OPTIONSTelemetry “from many popular libraries and frameworks without any code changes”
PythonStable, Stable, Developmentopentelemetry-bootstrap -a install, then opentelemetry-instrument python myapp.pyPopular libraries, “including Flask and Django”
PHPStable, Stable, StableThe opentelemetry extension plus the SDK and instrumentation packages through Composer, or the Linux-only PHP Distro packageWith the Slim and PSR-18 packages installed: a root span for the HTTP transaction, one for the action, one per outgoing HTTP call

Node.js APM with OpenTelemetry is one --require flag or one NODE_OPTIONS line. The Node.js zero-code page does not mention serverless functions; OpenTelemetry’s functions-as-a-service pages say “platform quirks usually mean these applications have slightly different monitoring guidance and requirements”, and the community currently provides “pre-built Lambda layers able to auto-instrument your application”.

Python performance monitoring works the same way: to monitor a Python application you install the opentelemetry-distro package, run opentelemetry-bootstrap -a install to add instrumentation libraries for the packages already installed, where one applies, and start the app under opentelemetry-instrument. The agent “primarily uses monkey patching to modify library functions at runtime”, which is why no code changes are needed.

To monitor a Java application with an open source tool, start with the OpenTelemetry agent from the table: application performance monitoring for Java then begins at JVM start-up with no code changes, and choosing among paid Java monitoring tools is left to the comparison page linked above. JVM profiling is a different question: Oracle’s Flight Recorder guide describes analyzing “in greater detail events generated by applications, the JVM, and the operating system”, which goes deeper than a first monitoring step needs.

APM for PHP starts with the opentelemetry extension, and OpenTelemetry’s PHP page warns that “Installing the OpenTelemetry extension by itself does not generate traces”: application performance monitoring on PHP also needs the SDK and an instrumentation package for your framework, on “PHP 8.0 or higher”. Route-level timing for each endpoint is what Laravel API monitoring asks for, and the Linux-only PHP Distro lists “Laravel 6.x to 13.x” among its instrumented frameworks.

Elastic APM is “an application performance monitoring system built on the Elastic Stack” that collects “response time for incoming requests, database queries, calls to caches, external HTTP requests, and more”, according to Elastic’s APM documentation. Elastic builds its own agents for Java, Node.js, Python, PHP, Go, Ruby and .NET, and recommends “using Elastic OpenTelemetry to collect application telemetry data”. If you self-manage the stack, Elastic’s APM Server “receives performance data from APM agents, validates and processes it, and transforms the data into Elasticsearch documents”, so you run that server and the Elastic Stack behind it. The hosting bill for that stack is yours, which is my reading of what self-hosting means for a two-person team.

Azure, GCP and cloud-function logging and monitoring

Google Cloud splits the job between two products: Cloud Logging stores, queries and can alert on logs, while Cloud Monitoring holds metrics, alerting policies and uptime checks. Azure Monitor is Microsoft’s umbrella, with Log Analytics for querying logs and Application Insights for application performance. Google keeps Data Access audit logs off by default, except for BigQuery.

GCP logging and monitoring best practices start with knowing which product holds what: in Google Cloud, logging sits in Cloud Logging and monitoring in Cloud Monitoring, both under Google Cloud Observability. Google’s Cloud Logging documentation calls Cloud Logging “a fully managed service that allows you to store, search, analyze, monitor, and alert on logging data and events from Google Cloud and Amazon Web Services”.

The _Default log bucket in a project keeps logs for 30 days by default, configurable from 1 to 3650 days; the _Required bucket keeps its logs 400 days and cannot be changed (Google’s quotas page, read 2026-09-28). Google’s pricing page lists the free allotment for log storage as the first 50 GiB per project per month (read 2026-09-28).

GCP audit logs come in four types. Cloud Audit Logs says Admin Activity and System Event logs are always written, Policy Denied logs are generated by default, and “Except for BigQuery, Data Access audit logs are disabled by default because they can generate large volumes of data.”

For monitoring and logging, Cloud Run functions start with what Cloud Run collects on its own. Google’s Cloud Run logging page says output to stdout or stderr is “picked up automatically by Cloud Logging”, and a single line of serialized JSON “is picked up and parsed by Cloud Logging and is placed into jsonPayload”, so each key becomes a field you can filter on.

Azure logging and monitoring best practices follow the same split under other names. Microsoft’s Azure Monitor overview calls Azure Monitor “Microsoft’s unified observability service”, says Log Analytics workspaces “collect log and trace data, which can be analyzed with Kusto Query Language (KQL)”, and describes Application Insights as offering “application performance monitoring (APM) for live web applications”. So Azure Monitor is the service, and Log Analytics is the workspace where its logs are stored and queried.

CloudWhere logs goWhere metrics and alerts liveDefault log retention (read 2026-09-28)The four things to turn on (my working rule)
Google CloudCloud Logging; Cloud Run picks up stdout and stderr on its ownCloud Monitoring: alerting policies and uptime checks_Default bucket 30 days, configurable; _Required bucket 400 days, not configurableAn error-rate alert; a log-based alert on one critical message; a budget alert on log volume; Data Access audit logs on the production database
AzureLog Analytics workspaces, queried with KQLAzure Monitor alerts; Azure Monitor workspaces for Prometheus and OpenTelemetry metrics; Application Insights for APMNot stated as a default in Microsoft’s overview; the pricing page says Analytics Logs can be retained 31 days at no chargeThe same four; for the database audit item, follow your database service’s own docs

The last column is my working rule, and Google attaches a condition to the audit item: “Enabling these logs might result in your Google Cloud project being charged for additional log usage.” That is why the budget alert goes in before the audit logs. If everything runs on one cloud, its own products can cover most of the setup table; Cloud Monitoring’s uptime checks, for example, let you “get notified when an endpoint fails to respond”. AWS logging and monitoring is a separate topic.

If the app calls a model: OpenAI API logs, Gemini API logs and LLM monitoring

LLM monitoring starts, by my working rule, with 7 fields logged for every model call: the model, latency, tokens in, tokens out, cost, the finish reason and the request id. Spend per hour is the first alert worth setting, then error and rate-limit responses, then latency.

OpenAI’s data controls guide says “By default, abuse monitoring logs are generated for all API feature usage and retained for up to 30 days, unless longer retention is required by law, or is reasonably necessary to protect our services or any third party from harm.” That retention exists for abuse monitoring. For debugging your own calls, keep your own record of each one.

Gemini API logs are viewed “in the Google AI Studio dashboard”, and you turn logging on or off per project from the Settings panel of the Logs and Datasets page. The Gemini API’s logs and datasets guide adds the condition: “Storage for Gemini API logs are only available for projects on the Gemini API paid tier.” A free-tier project has no stored logs to view, and stored logs expire after “a default retention window of 55 days” unless you save them to a dataset.

What to record yourself, per call, whatever the provider keeps, is my working rule in seven rows:

FieldWhy log itWhat never goes in the same line
ModelCost and behavior change from one model to the nextThe API key the call used
LatencyFeeds the third alert and shows a slow providerNothing to strip; log it as a number
Tokens inWith tokens out, the basis of the cost figureThe raw prompt when it holds personal data
Tokens outA sudden jump flags runaway answersThe raw answer when it repeats personal data
CostFeeds the spend-per-hour alertCard or billing details
Finish reasonSeparates complete answers from ones that stopped earlyNothing to strip
Request idJoins the call to the user’s request in your own logsThe user’s email or name in place of an id

For open-source LLM monitoring, Langfuse describes itself as “an open-source AI engineering platform” and its repository says it is “MIT licensed, except for the ee folders”. Self-hosting it still means paying for the servers it runs on, which is my reading rather than something its docs say.

Observability for startups: do you need tracing yet

Observability for a startup means 3 kinds of data joined by one request id: logs, metrics and traces. The request id in every log line and the basic metrics are needed from day one. Tracing earns its cost when there is a second service or a slow path nobody can explain.

Observability is being able to answer a new question about the system from the data it already emits, without shipping new code first. Observability logging best practices at this size reduce to one habit: every log line carries the request id, so logs, metrics and traces can be joined when the question comes. How logs, error tracking, traces and uptime differ is laid out in a table in the error logging article linked above, and the logging half of the work is the how-to-do-logging page, also linked above. Security events, such as failed logins and permission changes, are a separate list, covered in how can you prevent insufficient logging and monitoring.

How to check your own app

Application monitoring is verified by breaking 6 things on purpose: a test error, a failing health route, a crossed threshold, an alert read by someone who did not write it, a request id followed into the logs, and one dashboard number compared with its source. Record the date and result of each.

  1. 01 Throw a test error in staging and confirm it reaches the error tracker and the alert channel. Keep the date, the alert time and a screenshot of the alert. The full drill is in the error logging article.
  2. 02 Trigger a controlled check failure and verify notification and recovery reporting. Stop the health route or point the monitor at a path that fails, then restore it, and keep the time of the down alert and of the recovery message. The monitor's own test steps are in the uptime article.
  3. 03 Trigger each configured condition and record its alert threshold and behavior. Raise the error rate or slow a route in staging until each alert fires, wait out the provider's documented reporting delay before you read the result, and keep the threshold, the time it fired and the message.
  4. 04 Send test alerts and verify their destination, context, and response instructions. Then ask someone who did not write the alert to read one and say what they would do first, and keep their answer.
  5. 05 Take the request id from one error alert, search the logs for it, and find the lines that explain the error. Use the id your app writes on every log line and attaches to the error event: a host that prints its own request id holds a different value, so a search for that one finds nothing. Keep the query and the matching lines.
  6. 06 Compare the displayed metrics with their source systems and verify refresh behavior. Pick one dashboard number, such as the error count for today, check it against the tracker or the database, and note when it last refreshed.

Repeat all six after any change of host, error tracker or alert channel.

Where the sprint does this

In area 08 of the Production Hardening Sprint, deliverable 8.2 monitors the production URL and health endpoint with outage alerts ; 8.3 alerts on error spikes, latency, connection pressure, and queue backlog ; 8.4 routes alerts to the designated Slack or email destination and tunes thresholds to reduce noise ; and 8.8 creates one view of uptime, error rate, latency, signups, and revenue for the application’s relevant services. We verify each of the four the way checks 2, 3, 4 and 6 above describe. Hosting, paid tools, and API usage remain in your accounts, and we explain any required third-party costs before enabling them. Each deliverable and its verify line is listed in area 08 of the published scope.

Common questions about application monitoring

How do you monitor application performance?

Watch rate, errors and duration for each route, add an outside uptime check, and send alerts to one channel a named person reads. Add an APM agent when you need step-level timing, meaning which query or outbound call inside a request used the time.

Which APM tool is best?

No tool is best in general. The test is whether it covers your language with zero-code instrumentation, as OpenTelemetry’s agents do for Java, Python, Node.js and PHP, and what it costs at your request volume. The Sentry vs Datadog comparison, linked in the setup section, weighs named products.

What is the difference between Google Cloud Logging and Cloud Monitoring?

Cloud Logging stores, searches and alerts on log entries. Cloud Monitoring handles metrics, alerting policies on those metrics, and uptime checks that probe HTTP, HTTPS and TCP endpoints. Both sit under Google Cloud Observability, and a small app on Google Cloud usually needs both.

How long does Google Cloud logging retain logs?

The _Default log bucket in a project keeps logs 30 days unless you change it, and you can set 1 to 3650 days. The _Required bucket, which holds Admin Activity and System Event audit logs, keeps them 400 days and cannot be changed. These are the figures on Google’s quotas page as read on 2026-09-28.

Is Azure monitoring free?

Partly. Microsoft’s pricing page, read on 2026-09-28, says standard metrics and activity logs are collected at no cost, the first 5 GB a month per billing account of Analytics Logs ingestion on the Pay-As-You-Go tier is free, and Analytics Logs can be retained 31 days at no charge. Log data beyond that is billed by ingestion, retention and export.

What are the three types of metrics?

Prometheus names four core metric types, not three: a counter only goes up or resets to zero on restart, a gauge can go up and down, a histogram counts observations in configurable buckets, and a summary calculates configurable quantiles over a sliding time window. The Prometheus metric types page, linked in the RED section, covers each in depth.