The first line of any monitoring shopping list is an error tracker, and for most small apps that settles Sentry vs Datadog before the pricing pages open. Sentry finds the line of code that threw, and its Developer plan is $0 for one user. Datadog watches hosts, logs and metrics, billed per product, which matters once there are servers of your own to watch.

What we are comparing, and for whom: logging and monitoring tools for a small SaaS

This page is for a team of one to five people running an app with paying users on a managed host such as Vercel, Render, Railway or Supabase. You have no servers of your own to patch, and the question is which monitoring and logging tools to add, in what order, and what each one will cost as the app grows. The wider area, from signals to alerts, sits under logging and monitoring.

In my June and July 2026 audits, 17 of the 21 third-party apps had no error tracking or alerting: when a user hits an error, nothing records it. Those 21 are apps I audited and chose to look at, a selected set rather than a random sample, so read the figure as what I met, not as a rate for AI-built apps in general. My reading of it is that the first purchase is rarely the expensive one.

In one app I audited, an AI coding workspace had no error tracking in its code (the report notes the owner might run a log drain the code does not show), and an observability package was a declared dependency imported nowhere. The full telling, and the smallest setup these tools plug into, are in application monitoring best practices.

This is documented analysis, not a benchmark: every figure comes from the vendor’s own pricing page or documentation, and nothing here was load-tested or run side by side. Prices checked 28 September 2026.

For the monitoring tools used in DevOps, or a logging tools list, the table below gives one row per job rather than a top-ten ranking: the kind of tool that covers it, what a small SaaS gets for nothing, and the meter that ends the free tier. What each signal records and misses is already tabled in the error logging article linked in the next section, so this table adds only tool type, free allowance and meter.

The jobThe tool typeNamed examplesWhat a small SaaS gets freeThe meter that ends the free tier
ErrorsError trackerSentry, Rollbar, BugsnagSentry Developer: $0, one user, 5k errorsErrors per month, and a second user
Uptime from outsideUptime monitorSentry uptime monitoring, Better StackSentry Developer: 1 uptime monitor and 1 cron monitorNumber of monitors
LogsThe host’s log view, then a hosted log toolVercel runtime logs, Better Stack, DatadogVercel Pro keeps 1 day of runtime logs; Better Stack free keeps 3 GB of logs for 3 daysThe retention window, then gigabytes ingested
Metrics and dashboardsThe host’s panels, then a platformDatadog, New Relic, GrafanaDatadog Free: up to 5 hosts; New Relic: 100 GB of data ingest a monthHosts (Datadog) or gigabytes ingested (New Relic)
TracesAn APM agent or OpenTelemetrySentry tracing, Datadog APM, OpenTelemetrySentry Developer: 5M spansSpans per month, or APM hosts
Product analyticsA separate analytics toolNot compared hereNot compared hereNot compared here

Product analytics is a different purchase with its own shortlist, for the moment when you have no idea where users drop off. None of the rows is a ranking, and there is no single best pick among logging and monitoring solutions: whichever logging tool or software you pick, it fits only if its meter is the one you will hit last.

Sentry vs Datadog: the comparison, row by row

Sentry vs Datadog compares two tools built for different jobs. Sentry is an error tracker: it groups exceptions into issues with the stack trace, plus release and suspect commit once set up. Datadog is an observability platform: host metrics, logs, traces and dashboards, each billed as its own product. A small SaaS usually needs the error tracker first.

Read the table as Datadog vs Sentry on nine questions worth asking before signing up; each cell comes from that vendor’s own pricing page or docs.

QuestionSentryDatadog
The job it was built forFinding and fixing errors in application codeMonitoring infrastructure, applications and logs from one platform
Who its entry plan is written for”For solo devs working on small projects” (Developer plan)“Core collection and visualization features” (Free plan), with volume discounts from 500+ hosts a month
What an error looks like in itAn issue: grouped events with stack trace, plus release and suspect commits once configuredAn issue in Error Tracking, grouped from APM traces, RUM events or logs
Logs5GB included on every plan; $0.50/GB beyond on Team and Business$0.10 per ingested GB, plus $1.70 per million events indexed for 15 days
Infrastructure metricsHost metrics not stated in Sentry’s pricing; 5GB of application metrics includedInfrastructure Pro at $15 per host a month
Tracing and performance5M spans includedAPM from $31 per host a month with Infrastructure attached
Uptime and cron checks1 uptime monitor and 1 cron monitor on the free planSynthetic API tests at $5 per 10K test runs, a paid product
How it billsA plan fee with a quota per data type, then pay-as-you-go or reserved volumeEach product on its own meter: hosts, gigabytes, events, sessions, test runs
Free tierDeveloper: $0, one user, 5k errorsFree infrastructure tier with a host cap and 1-day metric retention

Datadog prices shown are the annual rates; on-demand rates are the same or higher. The two tools also connect, though Datadog frames the SDK route as a migration: Datadog’s page on Sentry events says you can point existing Sentry SDKs at a Datadog DSN “so you can start using Error Tracking on existing applications that are instrumented using Sentry SDKs”, and those events arrive in Datadog as logs. Datadog also documents a webhook route in the other direction, where a Sentry alert opens a page for a Datadog On-Call team.

Picking the tracker is the easy part. Wiring it, with releases and environments tagged, is the job of wiring an error tracker in one afternoon, and a tracker only pays off once its stack traces point at your source, which means fixing minified, unreadable stack traces.

What Sentry costs: the Developer plan, the Team plan and the quota a small app burns first

Sentry’s Developer plan is $0 for one user with 5k errors, and the Team plan is $26 a month when billed annually with default pre-paid data, with 50k errors. Errors are the quota a small app burns first, often from a single noisy exception, so set filters and a rate limit before upgrading.

The table reads straight from Sentry’s pricing page on sentry.io and its billing docs. Prices checked 28 September 2026.

PlanPriceUsersErrors includedWhat else is included
Developer$0One user5k errorsAlerts by email, 1 uptime monitor, 1 cron monitor, 5GB logs, 50 replays
Team$26/mo billed annually with default pre-paid dataUnlimited50k errorsAlerts through integrated tools, third-party integrations, a maximum spend threshold
Business$80/mo on the same annual basisUnlimited50k errorsAnomaly detection, advanced quota management, advanced inbound filtering, SAML and SCIM
EnterpriseCustomUnlimitedCustomTechnical account manager, dedicated support

Sentry is free for one person: the Developer plan is the Sentry free plan, and the account turns paid the moment a second person needs a login, because Developer is limited to one user. The Sentry free tier also keeps a shorter history than Team: a 30-day lookback against up to 90 days.

Sentry billing works per data type. Each plan includes a quota for errors, logs, spans, replays, attachments and monitors; extra volume comes either from reserved volume you pay for in advance at a discount, or from a pay-as-you-go budget “shared among all categories on a first-come, first-served basis”. Once both run out, Sentry drops further data and does not charge for it, which in its own words means “you’ll lose monitoring for the remainder of the billing cycle”.

Alert channels depend on the plan. Every Sentry plan sends alerts and notifications by email; alerts through integrated tools such as Slack are ticked for Team, Business and Enterprise and not for Developer.

Three prices get searched one by one. Sentry logs pricing comes first: 5GB is included on every plan, and more can only be bought from the pay-as-you-go budget at $0.50/GB. For replays, Sentry session replay pricing includes 50 replays, and on Team the next replays cost $0.0030 each reserved or $0.00375 each pay-as-you-go at the first volume tier. The AI debugger has its own Sentry Seer pricing: $40 per active contributor per month, billed separately from the pay-as-you-go budget, where an active contributor is anyone who makes 2 or more PRs to a Seer-enabled repository. With the plan table, those three cover what each of the Sentry plans costs.

Errors are the quota to watch, and in my reading a whole month of them usually goes in a burst: one exception thrown inside a loop, or a crawler hitting a broken page over and over. Sentry gives you three controls, and I’d set all three on day one. Inbound filters, in each project’s Inbound Filters settings, drop known noise such as web crawlers or a specific error message before it counts. A rate limit on the project’s client key, under Client Keys (DSN), caps how many error events that key accepts in a period. Spike protection, under Settings > Spike Protection, drops events once a project passes its spike threshold; it covers errors, spans and attachments, and it does not apply during a trial.

The free plan’s single uptime monitor and single cron monitor are enough to start watching one URL and one scheduled job. Whether you need more than that, and which checks matter, is covered in uptime monitoring for founders.

What Datadog costs: the free tier, per-host pricing and the log bill

Datadog’s pricing model has no single plan price: each product bills on its own meter. Infrastructure is per host per month, logs per gigabyte ingested plus per million events indexed, RUM per 1,000 sessions. The Datadog free tier is the Free plan of infrastructure monitoring: up to 5 hosts, with 1-day metric retention.

“Datadog plans” means per-product tiers, not one account plan. Infrastructure monitoring comes as Free, Pro and Enterprise, and the Datadog Pro plan is Infrastructure Pro at $15 per host per month billed annually, or $18 on demand. The table lists the main meters from Datadog’s pricing page. Prices checked 28 September 2026.

ProductThe meterList price, billed annuallyWhat the free tier includes
InfrastructurePer host per monthPro $15, Enterprise $23Free plan: 5 hosts at most, metrics kept for a day
APMPer host per month$31 with Infrastructure attachedFree trial only
LogsPer GB ingested, plus per million events indexed$0.10 per GB; $1.70 per million for 15-day retentionFree trial only
RUMPer 1,000 sessions$0.15Free trial only
Synthetics (API tests)Per 10K test runs$5Free trial only
Error TrackingPer error volume$25 a month for 50k errorsIncluded for APM traces and RUM events

Datadog is free to start: the Free infrastructure tier costs nothing, and every paid product card offers a free trial, which Datadog’s trial page sets at 14 days. Datadog infrastructure monitoring pricing then climbs per host, and on Pro each host carries 100 custom metrics. As for Datadog licensing, the pricing page lists annual, monthly and hourly plans; a perpetual license is not stated on Datadog’s pricing page.

A free alternative to Datadog comes in two forms: its own Free plan, which covers a few hosts with short retention, and the open source platforms further down, which cost nothing to download. Two billing mechanics, how hosts and custom metrics are counted, are explained in the FAQ on why Datadog is expensive. On a managed host with no servers of your own, the per-host meter has little to count, which is why the platform question can usually wait.

Datadog logs pricing: ingest, index and retention are billed separately

Datadog logs pricing has two meters: a charge per gigabyte ingested, and a charge per million log events indexed that rises with the retention window. A small app controls the bill by indexing only errors and warnings and archiving or dropping the rest.

Ingest is $0.10 per GB of uncompressed data. Indexing, billed annually per million events, is $1.06 for 3-day retention, $1.27 for 7 days, $1.70 for 15 days and $2.50 for 30 days, with longer windows priced on request. Archived logs can come back later: rehydrating costs $0.10 per compressed GB scanned, plus your indexing rate for the events it brings back.

The Datadog log cost you can control sits in the index. By default an index has no exclusion filter, so every matching log is indexed; an exclusion filter removes a slice from the index while it “still flow[s] through the Livetail and can be used to generate metrics and archived”, as Datadog’s log index docs put it. A daily quota puts a hard cap on how many logs one index stores per day. Together they keep Datadog log costs tied to the lines you read. Retention is its own decision, how long to keep application logs, and what belongs in a line is the error logging article’s section “Give every event a stable shape”. Datadog log management pricing beyond 30 days is by quote.

Datadog RUM pricing, serverless, database monitoring, SIEM and the other add-ons

Each add-on below has its own meter and price, billed annually, from Datadog’s pricing page on 28 September 2026. The last column is my working rule for an app with under a few thousand users.

Add-onWhat it isIts meter and list priceDoes a small SaaS need it yet?
RUMPerformance and errors from real browser and mobile sessions$0.15 per 1,000 sessions; Session Replay $2.50 per 1,000Later: once front-end complaints arrive
ServerlessMonitoring for functions and serverless apps$5 per active function a month; traced invocations $10 per millionLater: start with the host’s function logs
Database MonitoringQuery performance on database hosts$70 per database host a monthLater: first check the database dashboard
Product AnalyticsFunnels, retention and paths from sessions$0.80 per 1,000 sessionsNo: a separate analytics decision
Cloud SIEMSecurity detection over analyzed logs$5 per GB analyzedNo: built for a security team
Audit TrailA record of access and changes in Datadog itself2% of monthly spend, searchable up to 90 daysNo: only audits the Datadog account

RUM counts sessions on full traffic, so Datadog RUM pricing follows sessions, not signed-up users. Datadog serverless pricing counts, each hour, the functions that ran at least once, then averages that count over the month. For databases, Datadog database monitoring pricing counts database hosts, with 200 normalized queries allotted per host. Product analytics is also billed per session, so Datadog product analytics pricing tracks traffic the way RUM does. The Datadog Cloud SIEM pricing follows the gigabytes it analyzes, while Datadog audit trail pricing is a share of your monthly Datadog spend. SIEM and audit trail are security products; what a small SaaS should log for security is a separate question, covered in the series article on preventing insufficient logging and monitoring.

Datadog vs New Relic, Dynatrace, CloudWatch and Sumo Logic: the other platform matchups

Four more platform matchups share one answer shape: how the platform bills, what is free, and who its pricing is written for. The table puts them side by side; each H3 below takes one pairing. No winners, and every price is the vendor’s own, read on 28 September 2026.

PlatformHow it billsFree tierWho the pricing is written for
DatadogPer host and per productFree infrastructure tier: up to 5 hostsVolume discounts from 500+ hosts a month
New RelicPer GB ingested, plus user seats or compute100 GB of ingest a month and one full platform userAnyone: a perpetual free tier, then paid users by edition
DynatraceA minimum annual commitment drawn down at hourly rate-card pricesNot stated on the pricing page (a free trial is offered)Buyers who commit at the platform level
Amazon CloudWatchPer metric, per GB of logs, per alarm metric, per API requestMost AWS service metrics, plus small allowances for logs, custom metrics and alarmsAnyone on AWS, with no upfront commitment
Sumo LogicCredits, with a Flex option at $0 ingest and search billed by scanNone printed; a free trial insteadEssentials: small-to-medium DevOps and SecOps teams

Datadog vs New Relic: per host against per gigabyte

Datadog vs New Relic is a difference in meter. Datadog bills per host and per product. New Relic bills on data ingested plus either user seats or compute, with a free monthly data allowance. For one small app, compare each free allowance with your own log volume first.

New Relic’s free tier includes 100 GB of data ingest a month, one free full platform user and unlimited free basic users; beyond that, ingest is $0.40/GB on the original data option. Seats come in three types, and the first paid full user starts at $10 on the Standard edition, per New Relic’s pricing page. For one small app, that means the New Relic vs Datadog bill follows log volume and head count on one side, and hosts plus the number of products switched on on the other. Both accept OpenTelemetry data natively, so the instrumentation you add for one can later point at the other.

Datadog vs Dynatrace, and Dynatrace vs AppDynamics

Datadog vs Dynatrace is a choice between two enterprise platforms: Dynatrace sells one annual platform commitment drawn down at hourly rate-card prices, Datadog sells per-product meters. Dynatrace vs AppDynamics is an older enterprise APM matchup. Neither pairing is a small-team decision.

On Dynatrace’s pricing page, Infrastructure Monitoring is $29 a month per host, billed at $0.04 per hour, and Full-Stack Monitoring is $58 a month per 8 GiB host, billed per memory-GiB-hour, all drawn from “a minimum annual spend commitment at the platform level”. AppDynamics is sold today as Splunk AppDynamics, pitched at “hybrid and on-prem application performance”. The platform table above names the products Dynatrace’s pricing competes with, in no order.

If an enterprise customer’s security questionnaire asks which of these you run, read it as a question about whether you monitor production at all, with evidence, not as a demand for a particular brand.

CloudWatch vs Datadog

CloudWatch vs Datadog only matters when the app runs on AWS. CloudWatch already receives metrics from most AWS services automatically, billed per metric, per gigabyte of logs and per alarm metric, with a free tier. My working rule: use what the cloud gives you until it cannot answer a question.

The free tier covers 5 GB of log data, 10 custom metrics, 10 alarm metrics and 1 million API requests a month, with three API operations, GetMetricData among them, “always charged”. Datadog’s AWS integration reads those same metrics by a “metric-by-metric crawl of the CloudWatch API”, with new metrics pulled every ten minutes on average. That crawl is API traffic on your AWS account, and CloudWatch pricing charges API requests beyond the free million. Setting up the AWS side itself is a separate job: AWS logging and monitoring.

Sumo Logic vs Datadog, and the Sumo Logic price

Sumo Logic vs Datadog sets a log analytics and security product against a broader platform. The Sumo Logic price is sold as credits, with a Flex option that charges nothing for ingest and bills searches by the data they scan. For one small app the deciding number is daily log volume.

Sumo Logic’s pricing page offers two plans, Essentials and Enterprise Suite, paid in credits, and a free trial rather than a free tier. On Essentials, Sumo Logic pricing includes log retention of up to 365 days, and the page says costs vary with the subscription configuration you quote, so there is no single price to copy here. Datadog covers the same log ground plus infrastructure and APM. Before you compare either, read your daily log volume off the host’s own log view.

Sentry alternatives and the open source observability platform option

People searching for alternatives and people searching for open source are usually the same reader, one bill later. The first H3 covers tools that replace Sentry; the second covers what it costs to run the whole stack yourself.

Sentry alternatives, hosted and open source

Sentry alternatives fall into three groups: hosted error trackers such as Rollbar and Bugsnag, platforms with error tracking built in such as Datadog, New Relic and PostHog, and open source projects you host yourself, such as GlitchTip and SigNoz.

AlternativeTypeHosted or self-hostedFree tierThe trade
RollbarHosted error trackerHosted$0: 5K occurrences and 1K sessions a monthThe quota counts occurrences
BugsnagHosted error trackerHosted$0: 1 user, 7.5K events and 1M spans a month, 7 days of dataOne user, one week of history
Datadog Error TrackingPlatform featureHostedIncluded for APM traces and RUM eventsStandalone errors bill from $25 a month for 50k
New RelicPlatform with error trackingHosted100 GB of ingest a month, one full platform userError tracking needs a core or full user seat
PostHogPlatform with error trackingHostedSee its pricing pageNot compared here
GlitchTipOpen source error tracker that accepts Sentry SDKsBothHosted Free plan: up to 1,000 events a monthYou run it, or pay for hosting
SigNozOpen source observability platformBothSee its pricing pageA ClickHouse store to run and size

GlitchTip says its app “is compatible with Sentry client SDKs”, so a Sentry-instrumented app can switch by changing where the SDK sends. On Sentry logging and open source: Sentry’s SDKs are MIT licensed, while the Sentry server uses the FSL-1.1-Apache-2.0 license, which becomes plain Apache-2.0 after a two-year grace period. Sentry’s logs product itself is part of the hosted plans. In my reading, the usual reason to leave Sentry is a quota bill, and the cheaper fix is the three day-one controls in the Sentry cost section above.

What an open source observability platform costs to run

An open source observability platform is free to download and never free to run: it needs a server, storage that grows with every log line, upgrades and its own monitoring. Three options worth a look are the Grafana stack, SigNoz and OpenObserve, each listed by OpenTelemetry as taking its data natively.

The Grafana stack is Grafana for dashboards, Prometheus or Mimir for metrics, Loki for logs and Tempo for traces; Grafana, Loki and Tempo moved from Apache 2.0 to AGPLv3 after an announcement on April 20, 2021, and Prometheus is Apache 2.0. SigNoz is built on OpenTelemetry and stores data in ClickHouse, under an MIT Expat license outside its enterprise directories. OpenObserve is AGPL-3.0, written in Rust, and stores parquet files on local disk or object storage such as Amazon S3. All three take OpenTelemetry data, and OpenTelemetry calls itself “vendor- and tool-agnostic”, which is what makes a later switch cheap: see what OpenTelemetry is.

Open source application performance monitoring is the same recipe: OpenTelemetry’s agents in the app, one of those backends behind them. Top open source application monitoring tools are free to download, and their running cost is set out here as line items rather than a figure.

What you runWhat it storesThe monthly cost you still pay
A server or two for the storeMetrics, logs and tracesThe server, sized for the store, not for the app
Disk or object storageEvery log line you keepStorage that grows with volume and retention
Upgrades and patchesNothingSomeone’s hours every release
Backups of the storeA copy of the dataMore storage, plus a restore you have tested
Monitoring for the monitorIts own healthAn outside check, or a hosted free tier anyway

For a team of two, set those line items against a hosted free tier that costs nothing until volume grows, and the break-even is usually far away: the server alone is a monthly bill from the first day, while the hosted meters above start charging only at gigabytes or hosts a small app may not reach. That is my working rule, not a measured figure. Before you run open-source monitoring tools the DevOps way, name the person who will own the stack; without one, stay hosted. Tracing itself is a separate topic: OpenTelemetry distributed tracing.

Log tools: logging as a service, ELK, Splunk, Graylog, Loki and the shippers that feed them

The log-tool questions below all come down to one decision: where the log lines go, and who runs the store. The summary table answers it once; the H3s below take one family each.

ToolWhat it isLicenseStorage modelWho runs it
Better StackHosted log and uptime serviceCommercialHosted, retention by planThe vendor
ELK (Elasticsearch, Logstash, Kibana)Search engine, pipeline and dashboardElastic License 2.0, SSPL or AGPLv3 for Elasticsearch and KibanaSearch indexYou, self-managed
OpenSearchFork of Elasticsearch 7.10.2 with OpenSearch DashboardsApache 2.0Search index on Apache LuceneYou, self-managed
SplunkCommercial log and security analyticsCommercialNot stated on Splunk’s pricing pageSplunk Cloud, or you with Splunk Enterprise
Graylog OpenLog manager with its own search backendSSPLData Node or self-managed OpenSearchYou, or the paid Graylog Cloud Platform
Grafana LokiLog store indexed by labelsAGPLv3Label index, compressed chunks in object storageYou, or Grafana Cloud
Filebeat, Fluentd, Fluent BitShippers that collect and forwardFluentd: Apache 2.0None: they forwardYou, on each machine

Logging as a service, and Better Stack pricing as the small-team example

Logging as a service means shipping log lines to a hosted product that stores, searches and alerts on them, priced per gigabyte with a retention window. For a small SaaS, my working order is the host’s own log view first, then one hosted log tool on its free tier.

The same order holds for server logs: a log monitoring system you run yourself is the ELK, Graylog or Loki route below. When you compare application log monitoring tools, check five things on your own data: the price per gigabyte, retention on the free tier, how fast a search returns, whether it can alert on a log pattern, and whether you can export your logs. The best log monitoring tools for a small team are simply the ones that pass all five on a week of real traffic. Server log monitoring software that you host yourself adds the running costs from the open source section.

Better Stack pricing serves as the worked small-team example. Its free plan keeps 3 GB of logs for 3 days and includes 10 monitors and heartbeats with Slack and e-mail alerts; paid bundles keep logs for 30 days, starting with Nano at $25 a month billed yearly for 40 GB of logs, per Better Stack’s pricing page.

ELK vs Splunk, and the open source Splunk alternatives

ELK vs Splunk sets a stack you assemble and can run yourself against a commercial product you license. ELK is three projects: Elasticsearch, Logstash and Kibana. Splunk is priced by ingest, workload or activity. The open source Splunk alternatives are a short list: OpenSearch, Graylog and Loki.

In the ELK stack vs Splunk comparison, the license is the part that changed. Elasticsearch and Kibana moved from Apache 2.0 to a choice of SSPL and the Elastic License with the 7.11 release in 2021, and in September 2024 Elastic added AGPLv3 as a third option, per Elastic’s licensing FAQ. OpenSearch is derived from Elasticsearch 7.10.2, with OpenSearch Dashboards derived from Kibana 7.10.2, all under Apache 2.0. That license history is the part of the question about companies moving away from Elasticsearch that the sources answer; their reasons are their own.

On the Splunk side, Splunk’s pricing page offers activity-based, ingest or workload pricing in the cloud, and data-ingest or workload pricing for self-managed Splunk Enterprise. There is no single free alternative to Splunk to crown: an open source Splunk replacement, or an ELK stack alternative, is one of OpenSearch, Graylog or Loki, and the Kibana alternatives are OpenSearch Dashboards and Grafana. None of this is a small-SaaS purchase; in my reading, the reason to know the names is an enterprise customer’s questionnaire or a new hire’s habit.

Graylog open source, ELK vs Graylog and Graylog vs Loki

Graylog Open is the free edition of Graylog, a log manager whose recommended search backend is its own Data Node. ELK vs Graylog is assembling three projects against installing one product. Graylog vs Loki is a searchable backend against a label-only index, which keeps Loki’s index much smaller.

Graylog’s paid editions are Graylog Enterprise, Graylog Security, Graylog API Security and the Graylog Cloud Platform, and Graylog Open can also run on self-managed OpenSearch instead of Data Node. Graylog vs ELK, then, is mostly about who assembles the parts: ELK hands you three projects to wire, while Graylog installs as one product. Loki takes the other road: it “does not index the contents of the logs, but only indexes metadata about your logs as a set of labels”, so a query finds streams by label and then decompresses only the matching chunks. The open source Graylog alternatives are the names from the ELK section. Whether Graylog Open still counts as open source is answered in the FAQ.

Loki vs Prometheus, and ELK stack vs Grafana

Loki vs Prometheus is not a choice between rivals: Prometheus stores numeric metrics, Loki stores log lines using the same label model, and Grafana reads both. ELK stack vs Grafana is the real decision: a search stack with its own dashboard against a dashboard that reads many stores.

Prometheus’ overview describes a toolkit that stores metrics as time series, collects them “via a pull model over HTTP”, queries them with PromQL and hands alerts to an alertmanager. Loki’s overview describes a log store “inspired by Prometheus” that differs “by focusing on logs instead of metrics, and collecting logs via push, instead of pull”. Prometheus vs Loki, then, is metrics against logs, and a Loki and Prometheus pair under Grafana is the Grafana stack from the open source section.

In Kubernetes, Loki is the same Loki, deployed as microservices “designed to run natively within Kubernetes”, with an agent such as Grafana Alloy scraping pod logs, labeling them and pushing them in. For ELK stack vs Grafana, Kibana is built to “visualize, explore, and manage data in Elasticsearch”, while Grafana queries and visualizes data “no matter where it’s stored”. What the metric types are and how they are queried is covered in Prometheus metric types.

Log shippers: what Filebeat and Fluentd do

Log shippers are small agents that read log files or container output and forward the lines to a log store. The three you will meet: Filebeat is Elastic’s shipper, Fluentd and Fluent Bit are the CNCF ones. On a managed host with log drains you need none of them, because the drain does that job.

Filebeat is, in Filebeat’s overview, “a lightweight shipper for forwarding and centralizing log data” that watches the files you name and forwards events to Elasticsearch or Logstash. Fluentd’s docs call it “an open-source data collector for a unified logging layer”, Apache 2.0 licensed; it graduated in the CNCF on April 11, 2019, and Fluent Bit is its lightweight graduated sub-project for logs, metrics and traces. Fluentd logs can go to any of the stores above. The OpenTelemetry collector does the same job for all three signals, as “a vendor-agnostic way to receive, process and export telemetry data”.

On Vercel, drains forward logs to an outside service, but only on the Pro and Enterprise plans. Open source logging, then, is three parts: a shipper, a store and a viewer. The open source log monitoring tools in the table above cover the store and the viewer, and any open source logging and monitoring stack can take its lines from these shippers.

Zabbix vs Nagios vs Prometheus: server monitors a managed host rarely needs

Zabbix vs Nagios vs Prometheus compares three open source server monitors. Nagios runs checks through plugins, Zabbix adds agents, templates and graphs, Prometheus scrapes numeric metrics and alerts on them. All three watch machines you operate, and on a managed host the provider operates the machines.

ToolModelWhat it is built aroundLicense
Nagios CorePlugin checksPlugins, “thousands” of themGPL v2
ZabbixAgents reporting to a serverAgents, templates and graphsAGPLv3 from version 7.0
PrometheusPull-based metric scrapingTime series, PromQL and alert rulesApache 2.0

Zabbix’s manual calls it “an enterprise-class open source distributed monitoring solution”, and Zabbix’s license page puts every version from 7.0 under AGPLv3. Zabbix alternatives are the other two rows plus the hosted platforms above. Nagios vs Prometheus is checks against metrics: a Nagios check answers whether something is up, while Prometheus keeps the numbers over time. What replaced Nagios has no single answer; in my reading, Zabbix, Prometheus and the hosted platforms are the ones to evaluate, which is advice, not a count of who moved. If you run a VPS, this section is for you. Uptime from outside is a different control, covered by the uptime article linked in the Sentry cost section.

When to pick each: the free-tier stack, paid Sentry, or a platform

Picking a monitoring tool takes 3 questions, by my working rule: do errors reach a person today, does anything check the app from outside, and can you search yesterday’s logs. If any answer is no, the free-tier stack fixes it. Pay for Sentry when quota runs out, and for a platform when there are servers to watch.

Each option below is my working rule, written as fits and does not fit, with no winner.

The free-tier stack

An error tracker’s free plan, its included uptime monitor (listed in the Sentry plan table), the host’s log view and one alert channel. On Sentry’s free plan that channel is email; routing alerts to Slack comes with a paid plan, and the routing itself is a separate setup: Slack alerting. The host’s log view is short-lived: Vercel’s runtime logs docs keep 1 hour of logs on Hobby, 1 day on Pro, and 30 days only with Observability Plus, so an empty search there proves nothing about last week.

Fits a small team on a managed host with no compliance ask yet. Does not fit when a second person needs a tracker login (the Developer plan is limited to one user), when you need log search beyond the host’s window, or when a customer asks for retention evidence.

Fits when a second person needs a login, alerts must reach Slack, or the error quota is truly exhausted after the noise controls are set. Does not fit when the missing thing is logs or host metrics. Team, priced in the plan table above, is the step up. Business is worth it only for the features it adds over Team, among them anomaly detection, advanced quota management, advanced inbound filtering, SAML with SCIM, Code Owners and a BAA.

Datadog or another platform

Fits when there are servers, containers or several services to watch, someone owns the dashboards, and a single bill for everything is worth a premium. Does not fit when the app is a single deployment on a managed host, nobody has time to tune it, or the budget has no cap. Set a daily quota on each log index the day logs are switched on.

A self-hosted open source stack

Fits when data must stay on your own infrastructure, or someone on the team already runs it well. Does not fit when the monitoring would share a server with the app it watches, or when nobody will patch it. Whatever you choose, a single screen over it is a separate job: how to build an ops dashboard.

How to trial a tool before you pay for it

A monitoring tool is trialed, by my working rule, with 5 tests before money changes hands: a deliberate error arrives with a readable stack trace, the alert reaches someone on a channel the plan includes, a usage cap can be set, a day of data can be exported, and a second teammate’s cost is known.

Run them in this order, my working order, on the free tier or the trial, and keep the evidence each one names.

  1. 01 Send a deliberate error from the deployed app, through a test route or a flag you control that is behind authorization and switched off afterwards, and confirm it arrives as an issue with a readable stack trace and the right environment. Keep the issue link.
  2. 02 Confirm the alert for that issue reaches a place someone looks at night, on a channel the plan actually includes, and wait out the tool's reporting delay before calling it failed. Keep the alert timestamp.
  3. 03 Find the usage cap, rate limit or spike limit and set it. Keep a screenshot of the settings page.
  4. 04 Export one day of data, or confirm the API allows it. Keep the file.
  5. 05 Add a second teammate and read what the plan page says that will cost. Keep the quote.

Throw a real exception for the first test, not a request for a missing page: the point is whether an exception in your own code reaches the tracker. A readable trace from a production build needs source maps uploaded, which is the stack traces article linked in the comparison section, and the full drill for a pipeline you keep is the error logging article’s “Prove the pipeline with a failure drill”. Pass means all five leave their evidence; a tool that fails test one is not monitoring anything.

The Production Hardening Sprint verifies its error tracking deliverable, 6.3, the same way: send a test error and verify symbolication, environment attribution, and alert delivery.

Where the sprint fits

On this ground, the sprint’s deliverables are these: 6.3 installs error tracking such as Sentry with protected source maps, environment labels, and alerts; 8.2 monitors the production URL and health endpoint with outage alerts; 8.3 alerts on error spikes, latency, connection pressure, and queue backlog; 8.4 routes alerts to the designated Slack or email destination and tunes thresholds to reduce noise; 8.8 creates one view of uptime, error rate, latency, signups, and revenue for the application’s relevant services; and 13.1, the production readiness report, delivers the result for every scope item, the work completed, and its verification evidence. Hosting, paid tools and API usage are paid through your accounts, and we explain any required costs before enabling them. Each line is in the published scope.

Common questions about monitoring tool costs

Why is Datadog so expensive?

Two billing mechanics make a Datadog bill larger than a host count suggests. Hosts are metered hourly and billed on the high-water mark of the lower 99 percent of those hours: only the top 1 percent is dropped, so extra hosts that stay up longer than that set the month’s count. Custom metrics are counted per unique combination of metric name and tag values, host tag included, so a metric tagged with a user ID can become one custom metric per user.

Can Sentry be self-hosted?

Yes. Sentry publishes a minimal self-hosted setup “with no guarantees or dedicated support”, and its docs set the minimum at 4 CPU cores, 16 GB of RAM plus 16 GB of swap, and 20 GB of free disk. The license lets you run it for your own app but not sell it as a service. The cost is that server and the upkeep of every service on it, which the open source section above lists as line items.

Is Nagios still free?

Yes for Nagios Core, which Nagios describes as “completely free forever under GPL v2 license”. Nagios XI is the commercial product, though Nagios also offers a bundle that runs a free edition of XI in a virtual machine.

Is Graylog still open source?

Not by the Open Source Initiative’s definition. Graylog Open, the free edition, is published under the Server Side Public License (SSPL), though Graylog’s own site menu still files it under Open Source. Elastic’s licensing FAQ describes SSPL as a source-available license originally created by MongoDB, not approved by the Open Source Initiative.

Is Logstash still used?

Yes, Elastic still ships and documents it as “an open source data collection engine with real-time pipelining capabilities”. For simple forwarding it is optional: Filebeat can send straight to Elasticsearch, and an OpenTelemetry collector can do the forwarding for logs, metrics and traces.

What happened to Sumo Logic?

Sumo Logic was taken private. On May 12, 2023, Sumo Logic announced that Francisco Partners had completed its acquisition at $12.05 per share in cash, about $1.7 billion in equity value, and its stock stopped trading on NASDAQ.