The week an app moves from a managed host to its own server is the week somebody has to decide what to count. Prometheus metric types are the 4 answers on offer: a counter only goes up until a restart resets it, a gauge goes up and down, a histogram sorts observations into buckets, and a summary calculates quantiles in the app. Most small apps need the first three.

Prometheus metric types: counter, gauge, histogram and summary

Prometheus metric types are four: a counter is a total that only rises or resets, a gauge is a value that moves both ways, a histogram counts observations into buckets with a sum and a count, and a summary computes quantiles in the client. My working rule: count events with counters, measure levels with gauges, time things with histograms.

This page is one part of the logging and monitoring series, and the question of which signals deserve attention before any tool is chosen sits in application monitoring best practices. What follows is narrower: the four types, which one fits each thing you want to know, the names and labels to start with, and the queries that read them.

Prometheus’ metric types page defines them tightly. A counter is “a cumulative metric that represents a single monotonically increasing counter whose value can only increase or be reset to zero on restart”, used for requests served, tasks completed or errors. A gauge “can arbitrarily go up and down”: memory in use, open connections, the depth of a queue. A histogram counts observations, usually request durations or response sizes, into configurable buckets and also keeps a sum of all observed values. Its count comes as a series too. A summary samples the same kind of observations but calculates configurable quantiles over a sliding time window inside the app itself.

I settle gauge vs counter with one test: if the number can go down, it’s a gauge, and if the question you’d ask of it is “how many per second”, it’s a counter. Prometheus counters answer rates and totals; a Prometheus gauge answers “how much right now”.

TypeWhat it isAn example metric nameThe query that reads itThe common mistake
CounterA total that only increases or resets to zero on restarthttp_requests_totalrate(http_requests_total[5m])Using it for a value that can decrease
GaugeA single value that goes up and downqueue_depthavg_over_time(queue_depth[10m])Reading it with rate()
HistogramObservations counted into buckets, plus a sum and a counthttp_request_duration_secondshistogram_quantile(0.9, sum by (le) (rate(http_request_duration_seconds_bucket[10m])))More buckets or labels than the questions need
SummaryQuantiles calculated in the app over a sliding window, plus a sum and a countjob_duration_secondsjob_duration_seconds{quantile="0.95"}Averaging the quantiles of several instances

Each mistake in the last column has a documented reason. The docs say not to use a counter for a value that can decrease and to use a gauge instead. rate() “should only be used with counters”, and applied to a gauge it “will most likely produce a nonsensical result, but the query will be processed without complains”, so nothing warns you. Summary quantiles can’t be combined: the histograms guide says “you cannot aggregate quantiles” and marks an avg() over summary quantiles as “BAD!”.

Prometheus data types are a different list, the one the query language uses: an expression evaluates to an instant vector, a range vector, a scalar or a string, and the string type is “currently unused”. And outside Prometheus, types of metrics usually means business, product or performance metrics, which is a planning question rather than a storage one.

Prometheus histograms and how to choose histogram buckets

A Prometheus histogram counts observations into cumulative buckets, each labeled with its upper bound, plus a running sum and count. Percentiles are estimated from the buckets, so place bounds around the target you care about and keep the list short, because every bucket is another series for every label combination.

A classic histogram named http_request_duration_seconds exposes three kinds of series: http_request_duration_seconds_bucket with an le label for each upper bound, http_request_duration_seconds_sum and http_request_duration_seconds_count, and the count equals the le="+Inf" bucket. The buckets are cumulative: each one “counts all observations less than or equal to the upper boundary provided as a label”, so the 1-second bucket also holds everything in the 0.5-second bucket.

histogram_quantile() turns those buckets into an estimate. When the percentile you ask for falls between two bounds, the function “interpolates the quantile value within the bucket the quantile value falls into”, and for a classic histogram the error “is limited by the width of the bucket the quantile is located in”. That is the whole case for choosing buckets on purpose: a wide bucket around your target means a vague answer exactly where you need a sharp one.

Here is a Prometheus histogram example, worked from a target. Say the product rule is that a page should answer within 500 ms. My working rule is to put bounds at and around the target and keep the list to about ten, so for this target I’d start with 0.1, 0.25, 0.5, 1 and 2.5 seconds, written in seconds because the naming guide asks for base units. The Python client appends +Inf on its own, so that list gives six bucket series, plus _sum and _count, for each route that reports. With a bound sitting exactly on 0.5, the share of requests that met the target is an exact read rather than an interpolation:

sum(rate(http_request_duration_seconds_bucket{le="0.5"}[5m]))
/
sum(rate(http_request_duration_seconds_count[5m]))

The histograms guide shows the same pattern with a 0.3 bound and warns that “this expression strictly requires a bucket boundary configured at 0.3”; with no bound at your number, the classic form returns nothing at all. The Python client’s default buckets run from 0.005 to 10 seconds and “are intended to cover typical web/RPC request latency in seconds”, which is a fine start when you have no target yet.

Native histograms change this picture over time. The Prometheus guide to histograms and summaries now opens with “If you can, use native histograms and prefer them over both classic histograms and summaries”, and the native histogram spec says they are “supported as a stable feature” from v3.8.0, though scraping them “still has to be activated explicitly via the scrape_native_histograms configuration setting”. The same guide says native support in client libraries “is still rare” and needs the protobuf exposition format, naming the Java and Go libraries, so check what your own client supports before you plan on them. The math of percentiles, and why the slowest requests matter more than the average, is the subject of tail latencies.

A Prometheus metrics example: what a small SaaS should count and what it should time

A Prometheus metrics example for a small SaaS starts with a short list: request count and request duration by route, jobs processed and failed, queue depth, webhook outcomes, signups, database connections in use, and model spend if the app calls one. Label values stay a short fixed list, never a user id.

For a small SaaS I start from this list, eight rows that answer most of the questions a founder asks in the first month.

Metric nameTypeLabelsThe question it answers
http_requests_totalCountermethod, route, statusHow much traffic, and how much of it fails?
http_request_duration_secondsHistogramrouteWhich pages are slow, and is it getting worse?
jobs_processed_total, jobs_failed_totalCounterqueueAre background jobs finishing?
queue_depthGaugequeueIs the backlog draining or growing?
webhook_events_totalCounterprovider, outcomeAre payment and other webhooks landing?
signups_totalCounternoneDid signups stop after the last release?
llm_request_cost_dollars_totalCountermodelWhat are model calls costing, for apps that call one?
db_connections_in_useGaugenoneIs the app close to its connection limit?

The names follow Prometheus’ naming guide: base units such as seconds and bytes, the unit in the name, and _total on an accumulating count. Countable things like connections are not units under that rule, which is why db_connections_in_use carries no suffix. In Python, the client strips a _total you type into a counter’s name and adds it back when it exposes the series, so both spellings end up the same.

Labels need the most care. The guide’s caution is the one rule to remember: “every unique combination of key-value label pairs represents a new time series, which can dramatically increase the amount of data stored. Do not use labels to store dimensions with high cardinality (many different label values), such as user IDs, email addresses, or other unbounded sets of values.” My reading for web apps: the route label holds the route template, such as /invoices/:id, never the raw path, because a raw path carries an id and turns into exactly that unbounded set.

Picture a contractor adding Prometheus to a small SaaS and, to answer “which customers are busiest”, putting the user id on the request counter as a label. Every user who makes a request now creates new time series, one for each combination with the other labels, so the stored data grows with the user count instead of with the number of routes, which is what the naming guide cautions against. Label values have to be a short fixed list; per-user questions belong in the database or a log line. A label is a question you’ll ask of every sample, and a user id is not a question: it is a new series per user.

Why it matters for a small SaaS: numbers you own, on a server you run

17 of the 21 third-party apps had no error tracking or alerting: when a user hits an error, nothing records it. I audited those 21 apps in June and July 2026, and they are a selected set, not a random sample, so the figure is not a rate for AI-built apps in general.

An error tracker records the single event. Metrics add the trend: is checkout slower this week than last, is the job queue draining or piling up, did signups flatten after a release. That is my reading of where they earn their keep, and it’s why the starter list above mixes business counters with system ones.

Prometheus’ own overview says it “works well for recording any purely numeric time series” and fits “both machine-centric monitoring as well as monitoring of highly dynamic service-oriented architectures”. For a small SaaS I read that as a fit when the app runs on servers or containers you operate and you want signups and failed jobs on the same screen as memory and latency. When it doesn’t fit is in the questions at the end.

Prometheus and Grafana run on a server you pay for and patch, which is a real cost even when the software is free. One caveat decides where they run: a monitor on the same machine as the app goes down with it, which is the point made under whether a self-hosted monitor is enough. Keep the outside-in uptime check outside, and let Prometheus watch the inside.

Prometheus is metrics only: logging, traces and APM live elsewhere

Prometheus collects and stores metrics as time series, numbers with a timestamp and labels. Logs go to a log store such as Loki, traces to a tracing backend, and code-level APM needs a tracer. What Prometheus covers from your own instrumentation is the rate, errors and duration of every route.

The overview’s own description is that Prometheus “collects and stores its metrics as time series data”, with the timestamp and “optional key-value pairs called labels”. Prometheus logging has two honest readings. One is the log output of the Prometheus server itself. The other is where your app’s log lines go, and the answer is a log store running beside it, such as Grafana Loki. How the log tools compare is in Sentry vs Datadog and the other tool matchups, and what belongs in a log line is in how to do logging. For logs, error tracking, traces and uptime side by side, see logs, error tracking, traces and uptime told apart.

On Prometheus traces and Prometheus tracing: tracing is not stated in Prometheus’ overview, which describes metrics, rules and alerts. The one bridge is exemplars, an experimental feature of the exposition format: an exemplar “may have a Trace ID attached”, which lets a slow sample point at a trace kept in a tracing system. Tracing itself is a separate topic: OpenTelemetry distributed tracing.

Prometheus APM is the last misreading. Request rate, errors and duration from your own instrumentation cover part of what an APM tool shows; the per-request, code-level view of where the time went needs a tracer.

How it works: the endpoint, the scrape, the queries and the dashboard

Prometheus pulls. The overview lists “time series collection happens via a pull model over HTTP” among its main features, with a separate push gateway for short-lived jobs. For a small SaaS the whole loop has five parts:

  1. The app exposes a plain-text /metrics page through a client library.
  2. The Prometheus server scrapes that page, and every other target, on an interval and stores the samples.
  3. Rules run on the stored data: recording rules precompute expressions, alerting rules decide what is firing.
  4. Grafana sends PromQL queries to Prometheus and draws the answers.
  5. Alertmanager receives the alerts and handles “silencing, inhibition, aggregation and sending out notifications”.

The behavior below is taken from the Prometheus, Grafana and Docker documentation as it reads today. The compose file is a documented pattern, not a benchmarked production setup.

The Prometheus metrics endpoint, the default port and the Python client

A Prometheus metrics endpoint is a plain-text page, usually /metrics, that the server scrapes on an interval, and the server itself listens on port 9090 by default. Each metric has a HELP line and a TYPE line, then one line per label combination. The Python client builds it for you. Keep the endpoint private, because labels expose internals.

In the exposition format, lines starting with # are comments unless the next token is HELP or TYPE, and a TYPE line names the metric and one of counter, gauge, histogram, summary or untyped. Every sample line after that has to be a unique combination of metric name and labels.

On the Prometheus default port, the project’s port allocation list gives 9090 to the Prometheus server, 9091 to the Pushgateway and 9093 to Alertmanager, with exporters starting at 9100 for the Node exporter; cAdvisor sits on 8080. The Prometheus port is also where its own web UI and targets page live.

Your app’s endpoint should not be public. Metric names and label values describe routes, queues and providers, which is a map of your internals. Bind it to a private interface, keep it on the Docker network, or put it behind auth.

The Prometheus Python client installs with pip install prometheus-client. The Python client’s documentation covers a counter and a histogram like these:

from prometheus_client import Counter, Histogram, start_http_server

REQUESTS = Counter("http_requests_total", "HTTP requests", ["method", "route", "status"])
LATENCY = Histogram("http_request_duration_seconds", "Request duration", ["route"],
                    buckets=[0.1, 0.25, 0.5, 1, 2.5])

start_http_server(8000)  # serves the metrics page from a daemon thread on port 8000

def handle(method, route):
    with LATENCY.labels(route=route).time():
        status = "200"  # the real handler runs here
    REQUESTS.labels(method=method, route=route, status=status).inc()

start_http_server runs its own server “in a daemon thread on the given port”; inside an ASGI app such as FastAPI you mount make_asgi_app() at /metrics instead, and WSGI apps have make_wsgi_app(). Pre-fork servers such as Gunicorn need one more step. The client docs say the libraries “presume a threaded model, where metrics are shared across workers”, which “doesn’t work so well for languages such as Python where it’s common to have processes rather than threads”. Multiprocess mode needs the PROMETHEUS_MULTIPROC_DIR environment variable set to a directory the client library can use for metrics, plus a Gunicorn hook that marks dead workers. Without it, my reading is that each worker keeps its own counters and a scrape only sees the worker that answered. Node.js has the same shape: prometheus.io lists it among the client libraries, which it says “implement the Prometheus metric types”.

Prometheus and Grafana with Docker Compose, and monitoring Docker containers

Prometheus and Grafana run from one compose file with two services: the official Prometheus image with its config mounted, and the official Grafana Docker image, each with a named volume so data survives a restart. Pin the image tags, bind ports to localhost, and add cAdvisor when you want per-container metrics.

This is a Prometheus Grafana Docker Compose file for a single server, with the paths, ports and image names taken from the Prometheus installation page and Grafana’s Docker install guide. The tags are the latest releases on 2026-09-28; pin whatever is current when you set it up.

services:
  prometheus:
    image: prom/prometheus:v3.15.0
    volumes:
      - ./prometheus:/etc/prometheus:ro
      - prometheus-data:/prometheus
    ports:
      - "127.0.0.1:9090:9090"
    restart: unless-stopped
  grafana:
    image: grafana/grafana:13.2.2
    volumes:
      - grafana-storage:/var/lib/grafana
    ports:
      - "127.0.0.1:3000:3000"
    restart: unless-stopped
volumes:
  prometheus-data:
  grafana-storage:
# ./prometheus/prometheus.yml
global:
  scrape_interval: 15s
rule_files:
  - rules.yml
scrape_configs:
  - job_name: app
    static_configs:
      - targets: ["app:8000"]

The Prometheus image reads its config from /etc/prometheus and keeps its data in /prometheus, and the docs call a named volume “highly recommended” for production. The Grafana Docker image keeps its database in /var/lib/grafana and serves on port 3000; grafana/grafana is the open-source edition, and the guide says the default grafana/grafana-enterprise edition “is free and includes all the features of the OSS edition”. Compose puts every service on one default network where each container is “discoverable by its service name”, so the target app:8000 works when your app runs as a service called app in the same file, and in Grafana you add Prometheus as a data source at http://prometheus:9090.

The 127.0.0.1: prefix on each port line matters on a VPS. Docker publishes an unqualified port “to all host addresses”, and Docker’s note on published ports and ufw says that traffic to a published port “gets diverted before it goes through the ufw firewall settings”, so a firewall rule you think blocks port 3000 does not. With the localhost address, “only the Docker host can access the published container port”, and you reach Grafana through an SSH tunnel or a reverse proxy with auth. Grafana’s admin user and password both default to admin, and the password is “set once on first-run”, so change it the first time you sign in to the Grafana container.

Monitoring Docker containers themselves is a separate target. The cAdvisor guide runs cAdvisor, which “analyzes and exposes resource usage and performance data from running containers” in Prometheus format on port 8080, and reads per-container CPU, memory and network series such as container_memory_usage_bytes. The Docker engine can also be a target through metrics-addr in daemon.json, though Docker warns its metric names “may change at any time”. Either way, to monitor Docker with Prometheus is to add one more scrape_configs job. How long to keep all this on disk is a retention question, like how long to keep application logs, and an app on AWS is a separate topic: AWS logging and monitoring.

Prometheus rate, irate and why a counter is never read raw

Prometheus rate turns a counter into a per-second average over a window, such as five minutes, and copes with counter resets. irate uses only the last two samples, so it suits fast-moving graphs, not alerts. Apply rate first and sum second, never the other way round.

A raw counter is a running total since the process started, and it drops to zero whenever the process restarts, which includes most deploys, so its value alone answers nothing. rate(http_requests_total[5m]) gives “the per-second average rate of increase” over the five-minute window, and “breaks in monotonicity (such as counter resets due to target restarts) are automatically adjusted for”. PromQL’s functions reference says the Prometheus rate function “is best suited for alerting, and for graphing of slow-moving counters”.

On Prometheus rate vs irate, the docs are blunt: irate works from the last two data points and “should only be used when graphing volatile, fast-moving counters”, while rate is for alerts, because “brief changes in the rate can reset the FOR clause”. The PromQL rate rule that matters most is the order: “always take a rate() first, then aggregate. Otherwise rate() cannot detect counter resets when your target restarts.” So sum(rate(x[5m])) is right and rate(sum(x)[5m:]), a subquery that sums first, is not.

increase() is the same calculation told as a total: the docs call it “syntactic sugar for rate(v) multiplied by the number of seconds under the specified time range window”, meant “primarily for human readability”, and they want rate in recording rules.

Grafana rate is the same function typed into a Grafana panel, with one difference worth using. Grafana’s $__rate_interval variable picks the window for you and “guarantees a range window large enough to capture at least four scrape samples”, computed as max($__interval + scrape_interval, 4 * scrape_interval). It only works if the Prometheus data source’s scrape interval matches your real one; left at a default that’s too short, Grafana says it “causes rate() to return no data because there aren’t enough data points in the window”.

The Prometheus query language: a PromQL cheat sheet of ten queries

The Prometheus query language, PromQL, selects time series by name and label, then applies functions and aggregations. Ten queries cover a small SaaS: request rate, error ratio, p95 latency, traffic by route, the busiest routes, queue depth, failed jobs, window totals, a smoothed gauge and down targets.

The docs describe PromQL as “a functional query language” that “lets the user select and aggregate time series data in real time”. This PromQL cheat sheet runs against the starter metrics above, so each row is a query you can paste once those exist. Typed in order, the ten rows double as a short PromQL tutorial: selectors first, then rates, then aggregations and functions.

The questionThe queryWhat it returns
Requests per secondsum(rate(http_requests_total[5m]))One number for the whole app
Error ratio (5xx over all)sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m]))A fraction from 0 to 1
p95 latencyhistogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))Seconds, estimated from the buckets
Group by route: traffic per routesum by (route) (rate(http_requests_total[5m]))One series per route
The five busiest routestopk(5, sum by (route) (rate(http_requests_total[5m])))The top five route series
Queue depth nowqueue_depthThe latest value per queue
Jobs failed in the last hoursum(increase(jobs_failed_total[1h]))A count, possibly fractional
Total of a gauge’s samples over an hoursum_over_time(queue_depth[1h])The sum of every sample in the window
A smoothed gaugeavg_over_time(queue_depth[10m])The average per queue over ten minutes
Targets that are downup == 0Only the targets whose last scrape failed

Three rows need a note. Prometheus group by is the by clause on an aggregation: by “drops labels that are not listed”, and without removes the listed ones and keeps the rest. Prometheus sum over time adds up every sample in the window, which makes sense for a gauge you want to total or divide by count_over_time; on a counter the samples are already running totals, so summing them counts the same requests again and again, and increase() is the right tool. And up is 1 when a target was reachable and 0 when the scrape failed, while == 0 filters, dropping every series where the comparison is false. The increase row can print a fraction because the increase is extrapolated to cover the full window, so a counter that only grows in whole steps can still give a non-integer result.

These Prometheus query examples lean on four Prometheus functions, rate, increase, histogram_quantile and the _over_time family, plus the aggregation operators on PromQL’s operators page. A PromQL sum by and a PromQL topk are the two aggregations you’ll type most.

How PromQL differs from SQL, in my reading, comes down to two things. Every result is a set of labeled series at a point in time, not rows in a table. And there is no GROUP BY statement, only by and without on each aggregation, which is why a PromQL example often reads inside out.

Prometheus recording rules: precompute the query every panel repeats

Prometheus recording rules run a query on a schedule and store the answer as a new series, so an expensive expression is computed once instead of by every panel and alert. Name them level, metric, then operations, separated by colons, as the docs recommend.

The docs describe recording rules as a way to “precompute frequently needed or computationally expensive expressions and save their result as a new set of time series”, which they call “especially useful for dashboards”. Use them for the expressions every panel and alert repeats, such as the per-route request rate and the error ratio:

# ./prometheus/rules.yml
groups:
  - name: app
    rules:
      - record: job_route:http_requests:rate5m
        expr: sum without (instance, method, status) (rate(http_requests_total[5m]))
      - record: job:http_request_errors_per_requests:ratio_rate5m
        expr: sum without (instance, method, route, status) (rate(http_requests_total{status=~"5.."}[5m])) / sum without (instance, method, route, status) (rate(http_requests_total[5m]))

The names follow the level:metric:operations form from the rules practice page: _total is stripped when rate() is applied, a ratio joins its metrics with _per_, and every aggregation says what it drops with without. Check the file with promtool check rules before a reload. Alerting rules live in the same kind of rule group, so the error ratio you record here is what an alert can compare against a threshold; choosing the thresholds worth one is the question of how to alert on error rate spikes, and getting the message to a person is Slack alerting.

Grafana on top: template variables, a moving average, the dashboard API and iframes

Grafana template variables turn one dashboard into many by filling a dropdown from a label query. A moving average is avg_over_time in the query or a panel transformation. Dashboards are JSON behind an HTTP API, so keep them in the repository, and treat any iframe embed as a publishing decision.

Grafana’s variables docs describe variables “displayed as drop-down lists” at the top of a dashboard. For Prometheus, a query variable of type “Label values” with the label route and the metric http_requests_total fills the dropdown with your route templates, and panels then filter with route=~"$route"; the =~ is needed once multi-value or “All” is on. One dashboard now serves every route or queue.

A Grafana moving average has two homes. In the query it’s avg_over_time(queue_depth[10m]). In the panel it’s the “Add field from calculation” transformation in “Window functions” mode, whose “Mean” option “calculates the moving mean or running average”. My reading: prefer the query, because a transformation lives in one panel while an alert or recording rule reads the expression, and you want the graph and the alert to agree on the number.

The Grafana API for dashboards treats them as data. Dashboards “are represented as JSON objects”, and from Grafana 12 on the dashboard API exposes create and get endpoints under /apis/dashboard.grafana.app/v1/namespaces/:namespace/dashboards. That means you can export each dashboard, keep the JSON in the repository beside prometheus.yml, and restore it after a rebuild; Grafana’s dashboard API lists the calls.

A Grafana iframe needs two decisions. allow_embedding defaults to false, which sends X-Frame-Options: deny so browsers refuse to render Grafana in an iframe, a guard against clickjacking. Turning it on for an internal page is fine; pairing it with anonymous access under [auth.anonymous] so the frame loads without a login publishes your metrics to anyone who finds the URL, which is a leak unless the dashboard is meant to be public. What belongs on the one screen a founder reads is a separate question: how to build an ops dashboard.

Grafana and OpenTelemetry: where the collector fits

Grafana and OpenTelemetry meet through a store, not directly: the OpenTelemetry Collector receives metrics, logs and traces from the app and forwards the metrics to Prometheus for Grafana to query. An app that already exposes a metrics endpoint can skip the collector until it wants traces.

OpenTelemetry is a way to instrument code; Grafana is a viewer that reads telemetry from a store. The OpenTelemetry Collector “offers a vendor-agnostic implementation of how to receive, process and export telemetry data”. On the Prometheus side, Prometheus’ OpenTelemetry guide says the OTLP receiver is off by default “because Prometheus can work without any authentication”, and the flag --web.enable-otlp-receiver turns it on at /api/v1/otlp/v1/metrics.

Grafana’s own OTel collector distribution is Grafana Alloy, which its docs describe as a distribution of the OpenTelemetry Collector with extra integrations for the Grafana ecosystem. For a small SaaS already exposing /metrics, the collector is optional until you want traces, and the OpenTelemetry docs themselves say “in a development or small-scale environment you can get decent results without a collector”. Whether to adopt OpenTelemetry at all is the tracing page’s call, linked above.

How to check your own app: seven checks that the numbers are real

A Prometheus setup is proven by checks that a broken setup fails: every target shows UP, the metric appears with its TYPE line, ten test requests raise the counter by exactly ten, the rate query turns non-zero, no label grows without bound, and a deliberate failure drives an alert to firing.

These checks assume a Python app using the official client, scraped by the Prometheus server from the compose file above, on a staging or local instance with no other traffic. On Gunicorn or another pre-fork server, the client’s multiprocess mode has to be configured, or check 3 fails, which is the point of it. Write the date beside each result.

  1. 01 Open the Prometheus web UI, go to Status then Target health (the /targets page on port 9090), and confirm every target shows UP. Save a screenshot of your own page.
  2. 02 Request /metrics from the app and find your metric with its # TYPE line. Save the output.
  3. 03 Read one route counter on /metrics, send ten requests to that route, and read it again: it rose by exactly ten. A different number means the route label is not the template, the requests reached another worker, or other traffic reached the instance.
  4. 04 Run the rate() query for that route while test traffic runs and expect a non-zero value after a few scrape intervals. An empty result means the window holds too few samples.
  5. 05 Count the series per route with count(count by (route) (http_requests_total)) and confirm it is no higher than the number of route templates and stays put when you request new ids. Save the count.
  6. 06 Make one alert condition true on purpose in staging and watch the rule reach firing on the Alerts page, passing through pending if it has a for duration. A rule that never shows as pending or firing fails.
  7. 07 Compare one Grafana number with its source, such as the signups count for today on the panel against a count in the database for the same window. Save both numbers and the time.

Check 1 reads the page the getting-started guide names, where each target shows UP or DOWN. Check 6 follows Prometheus alerting rules: elements that are “active, but not firing yet, are in the pending state”, and a rule without a for clause becomes active on the first evaluation. The wait in check 4 is there because Grafana’s docs note that rate() returns no data when there aren’t enough data points in the window.

On the Production Hardening Sprint, we verify deliverable 8.3 by triggering each configured condition and recording its alert threshold and behavior, and deliverable 8.8 by comparing the displayed metrics with their source systems and verifying refresh behavior.

Where the sprint fits

Four deliverables in area 8 of the published scope cover this ground: 8.2 monitors the production URL and health endpoint with outage alerts ; 8.3 alerts on error spikes, latency, connection pressure and queue backlog ; 8.4 routes alerts to the designated Slack or email destination and tunes thresholds to reduce noise ; and 8.8 creates one view of uptime, error rate, latency, signups and revenue for the application’s relevant services . We refactor or replace components where the production work requires it, and your app’s current framework and hosting setup are our starting point. Hosting, paid tools and API usage are paid through your accounts, and we explain any required costs before enabling them.

Common questions about Prometheus and Grafana

When not to use Prometheus?

Prometheus’ overview gives its own answer: “If you need 100% accuracy, such as for per-request billing, Prometheus is not a good choice, as the collected data will likely not be detailed and complete enough.” Two more readings of mine for a small SaaS: skip it while a managed host’s built-in metrics already answer your questions, and skip it when what you need is logs or traces rather than numbers.

What is Prometheus vs Grafana?

Prometheus scrapes, stores and queries metrics, and runs the rules that record new series or raise alerts; Grafana draws dashboards from data sources such as Prometheus. The overview puts the split plainly: “Grafana or other API consumers can be used to visualize the collected data.” The compose file above runs the two side by side.

Can I use Prometheus for free?

Yes. Prometheus is released under the Apache License 2.0 and is a CNCF project that reached the Graduated maturity level on August 9, 2018. Grafana’s open-source edition is licensed under AGPLv3, and its Docker guide says the Enterprise image “is free” too. The real cost is the server they run on and the hours spent patching and upgrading it.

Can you push metrics to Prometheus?

Not by default: Prometheus pulls, and pushing goes “via an intermediary gateway”. The Pushgateway page is direct: “Usually, the only valid use case for the Pushgateway is for capturing the outcome of a service-level batch job.” A pushed job also loses the automatic up health series. The other route is the OTLP receiver, which accepts pushed OpenTelemetry metrics once it’s enabled with --web.enable-otlp-receiver.

Does Prometheus do alerting?

Half of it. Alerting rules are evaluated inside the Prometheus server, which sends the resulting alerts to Alertmanager, and Alertmanager “manages those alerts, including silencing, inhibition, aggregation and sending out notifications via methods such as email, on-call notification systems, and chat platforms”. Without Alertmanager, a firing rule shows on the Alerts page and notifies nobody.

What is the key difference between a summary and a histogram in Prometheus?

Where the percentile is computed. A summary calculates its quantiles inside your app and exposes them ready-made, while a histogram exposes bucket counts and the Prometheus server estimates the percentile with histogram_quantile(). That’s why only the histogram can be combined across several instances; the docs list summaries as “Not aggregatable”.