OpenTelemetry’s homepage lists a complete toolkit: APIs, SDKs, a Collector and a wire protocol. It never says when an app is too small to need it. In my judgment, OpenTelemetry distributed tracing earns its setup cost once a request crosses 2 or more services you run, or a slow path cannot be explained from logs. Before that, a request id on every log line does the job.
What OpenTelemetry distributed tracing is: spans, traces and one open standard
OpenTelemetry distributed tracing follows one request across every service it touches. OTel works by recording each step as a span with a name, start, end and attributes, and the spans share a trace id passed between services in the traceparent header. OpenTelemetry is the CNCF’s vendor-neutral standard for producing that data; storing and viewing it is a separate choice.
Traces are one signal among the several that make up logging and monitoring for a small SaaS, and usually the last one a small team adds.
In OpenTelemetry’s own definition, it is a CNCF project and “an observability framework and toolkit” for generating, exporting and collecting telemetry such as traces, metrics and logs. It says plainly that it “is not an observability backend itself”: the storage and the visualization are “intentionally left to other tools”.
A span is one unit of work. In OpenTelemetry’s traces concepts, each span carries a name, a parent span id (empty for the root span), start and end timestamps, attributes, events, links and a status. OTel spans nest: a child span is a sub-operation of its parent, and every span in one request shares the same trace id, which is what makes them a trace.
The “distributed” part is context propagation. When service A calls service B, A sends its trace id and span id along with the request, and B starts a new span in the same trace with A’s span as its parent. OpenTelemetry’s default propagator uses the W3C Trace Context headers, the traceparent header among them, and instrumentation libraries usually handle this for you. I’d draw the line here: timing work inside one process is a profiler’s or a logger’s job, and tracing becomes distributed the moment the request leaves that process.
The five words you will meet first are below. The “what it is” column is in OpenTelemetry’s terms; the “where it lives” column is my sketch of a typical small app.
| Word | What it is | Where it lives |
|---|---|---|
| Span | One unit of work or operation, with a name, timestamps, attributes, a parent and a status | Created inside your app, one per request handler, query or outbound call |
| Trace | The spans that share one trace id, arranged as a tree by their parent ids | Assembled in the backend you send spans to |
| Context propagation | Passing the trace id and span id from caller to callee so the callee’s span joins the same trace | In request headers between your services (traceparent) |
| SDK | The language library that implements the specification and APIs and exports the telemetry | In each service’s process, configured at start-up |
| Collector | A proxy that receives, processes and exports telemetry | A separate service, optional at small scale |
What it replaces: a vendor’s own instrumentation format, so the code that produces traces no longer decides where they go, and two older projects, OpenTracing and OpenCensus, which merged to form it. What it does not replace: a place to store and read the traces, an error tracker that groups exceptions, or your logs.
What is OTLP? The OTel protocol and the data format
OTLP, the OpenTelemetry Protocol, is how traces, metrics and logs travel between an SDK, a Collector and a backend. It runs over two transports, gRPC and HTTP. A backend that accepts OTLP can replace another by changing an endpoint setting rather than application code, which is what vendor-neutral means in practice.
The OTLP specification was at version 1.11.0 on 3 October 2026, stable for the trace, metric and log signals. Its default ports are 4317 for OTLP/gRPC and 4318 for OTLP/HTTP, and OTLP/HTTP sends Protobuf payloads in either binary or JSON encoding, with trace data going to the /v1/traces path by default.
OTel is the project; the OTel protocol is its wire format, the one part every SDK, Collector and compatible backend agrees on. So when a vendor says it accepts the OpenTelemetry format, it usually means OTLP. OTLP data is the spans, metric points and log records plus a resource, which OpenTelemetry defines as “the entity producing telemetry” described in resource attributes, such as service.name, “the logical name of the service”.
What is OpenTracing? OpenMetrics vs OpenTelemetry
OpenTracing was an earlier vendor-neutral tracing API, and OpenCensus a parallel project; the two merged to form OpenTelemetry. The OpenTracing project’s own site now says it “is archived”. OpenTelemetry’s OpenTracing migration page offers an OpenTracing shim in each SDK so old calls keep working during a move, and warns that “many OpenTracing libraries will be retired and may no longer be updated”. A tutorial that still imports OpenTracing is out of date.
OpenMetrics vs OpenTelemetry is a format against a framework, not two rivals. OpenMetrics describes itself as a specification “built upon and carefully extending Prometheus exposition format”, and the OpenTelemetry specification has a compatibility section that maps Prometheus and OpenMetrics metric points to OTLP and back.
An OpenTelemetry histogram records a population of measurements in a compressed form: a count, a sum and buckets, either with explicit boundary values or as an exponential histogram whose bucket boundaries follow an exponential formula. OpenTelemetry’s metrics page gives request latencies as its example of what a histogram aggregates. Bucket choices and the queries that read them belong with Prometheus metric types.
Why it matters for a small SaaS: tracing solutions, and whether you need one yet
Tracing solutions pay for themselves once a request crosses two or more services you run, a queue or a chain of model calls, or when a slow path cannot be explained from logs. An app with one service and one database is served by a request id on every log line and the database’s own slow query view.
That line is my own, not a standard. It agrees with the line in error logging best practices: “A single-service app can begin with request_id. Add trace context when background jobs, functions, queues, or external calls make one interaction difficult to follow.” How the id itself is generated and carried between services is the subject of trace ids and correlation ids.
Here is how I sort the signs:
| Sign in your app | What it means | Tracing or a request id |
|---|---|---|
| One deployable service and one database | Every request starts and ends in one process | Request id |
| Slow pages explained by the database’s slow query view | The time is in a query you can already see | Request id |
| Few outbound calls, each logged with its duration | The logs already show where the time goes | Request id |
| Two or more of your own services handle one request | Each service’s logs show only its own part | Tracing |
| A queue or background job picks up work from a request | The work finishes after the request has returned | Tracing |
| A chain of external or model calls per request | The slow link changes from one request to the next | Tracing |
| A slow path nobody can explain from logs | The missing time sits between log lines | Tracing |
| A customer asks where the time went on one request | You need that one request’s timeline | Tracing |
Before any of that, an app needs to know when it fails at all. Of the 21 third-party apps in my June and July 2026 audits, 17 had no error tracking or alerting: when a user hits an error, nothing records it. I chose the 21 apps I audited, so the count speaks for them and not for AI-built apps in general. My working order follows from it: an error tracker and a request id first, tracing after.
Take a founder whose checkout is slow for some customers: the Next.js front end on Vercel calls the founder’s own Python API, which calls the database and a payment provider, and the logs on each side show the request arriving and leaving but not where the time goes. With @vercel/otel on the front end only, the trace shows one long span for the call to the API and nothing inside it; once the API is instrumented and reads the incoming trace header, and the front end is set to send that header to the API’s domain, the same trace shows the database query and the payment call as child spans, so the slow step has a name. Tracing pays when a request crosses services you run, only across the hops that carry the context, and the span with nothing inside it names the service to instrument next.
Tracing solutions come in three shapes, as I group them: a vendor’s own agent and backend; OpenTelemetry instrumentation sending to a vendor’s backend; and OpenTelemetry sending to a backend you host yourself. The second and third keep the instrumentation the same and change only where the data goes.
OpenTelemetry vs Datadog is a category question more than a product one. OpenTelemetry is a standard and toolkit for producing telemetry; Datadog is a hosted platform that ships its own Agent and also takes OpenTelemetry data, through its Agent, through an OpenTelemetry Collector exporting over OTLP, or by direct OTLP ingestion. Datadog’s own docs add that pairing the OpenTelemetry API with Datadog’s SDK unlocks more Datadog features than the OpenTelemetry SDK alone. The vendor choice itself, Sentry vs Datadog and what each costs included, is a separate decision from adopting the standard.
How it works: instrumentation, the Collector, and somewhere to look at traces
The five parts below come in the order a team meets them. The smallest working path is the SDK with zero-code instrumentation, exporting OTLP straight to a backend, with no Collector in between. OpenTelemetry’s own Collector docs call sending data directly to a backend “a great way to get value quickly” when getting started, and say a small-scale environment “can get decent results without a collector”.
OTel environment variables and the shortest OpenTelemetry tutorial
OTel environment variables are defined once in the OpenTelemetry specification, so the language SDKs that support them read the same names. I’d start with three: OTEL_SERVICE_NAME, OTEL_EXPORTER_OTLP_ENDPOINT, and OTEL_EXPORTER_OTLP_HEADERS for the backend’s key. OTEL_TRACES_SAMPLER and its argument set how many traces are kept.
Support is optional: the SDK environment variable spec says implementations “MAY choose to allow configuration via the environment variables”, and those that do should use these names, so check your language’s SDK docs before relying on one.
| Variable | What it sets | Example value |
|---|---|---|
OTEL_SERVICE_NAME | The service.name resource attribute; wins over a service.name set in OTEL_RESOURCE_ATTRIBUTES | checkout-api |
OTEL_EXPORTER_OTLP_ENDPOINT | The base URL for every signal; over OTLP/HTTP, traces go to v1/traces under it. Default http://localhost:4318 for OTLP/HTTP, http://localhost:4317 for OTLP/gRPC | http://collector:4318 |
OTEL_EXPORTER_OTLP_HEADERS | Key-value pairs sent with each export, in the W3C Baggage format | key1=value1,key2=value2 |
OTEL_EXPORTER_OTLP_PROTOCOL | grpc, http/protobuf or http/json; the spec’s preferred default is http/protobuf, though an SDK may keep grpc | http/protobuf |
OTEL_TRACES_SAMPLER | The sampler; default parentbased_always_on | parentbased_traceidratio |
OTEL_TRACES_SAMPLER_ARG | For the ratio samplers, a probability from 0 to 1; default 1.0 | 0.25 |
OTEL_RESOURCE_ATTRIBUTES | Key-value pairs used as resource attributes | deployment.environment.name=staging |
The shortest working tutorial, following OpenTelemetry’s Python zero-code docs, has five steps. Install the distro and the OTLP exporter with pip install opentelemetry-distro opentelemetry-exporter-otlp, then run opentelemetry-bootstrap -a install to add instrumentation for the libraries you already use. Set OTEL_SERVICE_NAME and the endpoint. Start the app under opentelemetry-instrument. Make one request, and find its trace in the backend. Match the port to the protocol: 4317 for grpc, 4318 for http/protobuf.
The backend’s API key goes in OTEL_EXPORTER_OTLP_HEADERS, and I keep it in the host’s secret store, never committed to the repository. The per-language list of zero-code agents is part of application monitoring best practices and is not repeated here.
Instrumentation by stack: FastAPI, NestJS, Spring Boot, Go, PHP, Serilog, React and nginx
OpenTelemetry instrumentation differs by runtime. Python, Node.js, Java and PHP have zero-code options: a package, an agent or an extension that traces common libraries without code changes. For FastAPI, one package and one call trace incoming requests, and the database driver and the HTTP client need their own instrumentation packages.
The package and mode cells come from each project’s own docs, read on 3 October 2026; the last column is my pick of the one step that is easy to miss.
| Stack | Package or module | Zero-code or manual | One thing to add by hand |
|---|---|---|---|
| FastAPI | opentelemetry-instrumentation-fastapi, called with FastAPIInstrumentor.instrument_app(app) | Either: the call, or the opentelemetry-instrument launcher | The SQLAlchemy and httpx instrumentations, or queries and outbound calls stay invisible |
| NestJS | @opentelemetry/sdk-node with @opentelemetry/auto-instrumentations-node, which includes instrumentation-nestjs-core | Zero-code with --require @opentelemetry/auto-instrumentations-node/register | Load it before the app code, as the Node docs require |
| Spring Boot | spring-boot-starter-opentelemetry (Micrometer Tracing over OTLP), or OpenTelemetry’s Java agent or its own Spring Boot starter | The agent is zero-code; the starters are configured through Spring | On Spring’s starter, raise the sampling probability for testing; Spring Boot samples 10% by default |
| Go | Contrib libraries such as otelhttp for net/http | Manual; zero-code is still work in progress | Wrap each handler and database client yourself |
| PHP | Composer packages plus the PECL extension, or the Linux-only PHP Distro packages | Zero-code: PHP 8.0 or later with the extension, 8.1 to 8.4 with the Distro | Install the extension on the host, not only in Composer |
| .NET with Serilog | Serilog.Sinks.OpenTelemetry | Logs only, sent over OTLP | Keep the trace and span id fields on (on by default) |
| React (browser) | The OpenTelemetry web SDK | Manual; experimental | Decide whether to trace the browser at all |
| nginx | ngx_otel_module, prebuilt as nginx-module-otel since 1.25.3 | Configuration directives | Set otel_trace_context to propagate or inject; the default is ignore |
For FastAPI, the FastAPI instrumentation docs describe the package as instrumenting “http requests served by applications utilizing the framework”, which is why database and HTTP client spans need the SQLAlchemy and httpx packages beside it. The manual form of OpenTelemetry instrumentation for FastAPI is this:
import fastapi
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
app = fastapi.FastAPI()
FastAPIInstrumentor.instrument_app(app)
A FastAPI OpenTelemetry setup needs an SDK and an exporter before that call; the opentelemetry-distro package installs the API, the SDK and the opentelemetry-instrument launcher that configures them from the environment variables above.
For NestJS, the OTel setup is the Node.js SDK plus the auto-instrumentations package, and OpenTelemetry’s Node docs say the instrumentation “must be run before your application code”, either through --import with your own setup file or through the --require register hook. The NestJS OpenTelemetry instrumentation is in the package’s default list.
Spring Boot tracing has a documented path from each project, and each project points at its own. Spring Boot’s tracing chapter auto-configures Micrometer Tracing and needs spring-boot-starter-opentelemetry to report over OTLP; OpenTelemetry’s own Spring Boot page calls its Java agent “the default choice” and offers its own OpenTelemetry Spring Boot starter as the other option. On Spring’s own path, a Spring Boot trace from a test request can go missing, because by default Spring Boot “samples only 10% of requests” until you set management.tracing.sampling.probability=1.
Go has stable tracing and metrics SDKs, and its contrib instrumentation library for net/http “automatically creates spans and metrics based on the HTTP requests”; the Go zero-code project is described as “work in progress”. A Golang trace in the runtime sense is a different tool: the runtime/trace package feeds the Go execution tracer, which records goroutine, syscall and garbage-collection events inside one program, not requests across services.
OpenTelemetry PHP has stable traces, metrics and logs, and two zero-code routes: Composer with a PECL extension, which needs PHP 8.0 or later, or the PHP Distro packages, which run on Linux only.
For .NET teams on Serilog, OpenTelemetry arrives through the Serilog OpenTelemetry sink, which turns Serilog events into OpenTelemetry log records, sends them to an OTLP endpoint over gRPC or HTTP, and fills the trace id and span id from the current .NET Activity by default. That links logs to traces; the spans themselves still come from .NET’s own tracing.
OpenTelemetry for React and other browser code is the least settled part: the project’s browser guide says client instrumentation “is experimental and mostly unspecified”. For most small teams I’d stop at the first server hop and leave the browser out until a real question needs it.
nginx’s OTel module docs cover the proxy hop: ngx_otel_module supports W3C context propagation and the OTLP/gRPC export protocol, is available as the prebuilt nginx-module-otel package since 1.25.3, and is loaded as a dynamic module with load_module. Tracing is off until otel_trace on, and the header is left alone until otel_trace_context says otherwise.
On every stack, I add one attribute by hand: the tenant or a hashed user id, so you can find one customer’s traces; the email address and the prompt text never go into span attributes.
Vercel OTel: tracing a Next.js app on Vercel
Vercel OTel is the @vercel/otel package plus an instrumentation file that Next.js runs at start-up. It registers the OpenTelemetry SDK with fetch instrumentation and the W3C Trace Context propagator, and exports through a Vercel tracing integration if one is configured, otherwise through an OTLP exporter set in environment variables.
Vercel’s tracing docs give the set-up: install @opentelemetry/api and @vercel/otel, then add instrumentation.ts at the project root (or in src if you use one) with a register() function that calls registerOTel({ serviceName: 'your-project-name' }). Next.js then contributes a root span for each incoming request and a fetch span for each fetch your code makes.
The default that catches people: @vercel/otel propagates the trace context only to the deployment’s own URLs. A fetch to an API on another domain shows as one client span with its duration, and nothing that happens inside the API joins the trace until that domain is listed in propagateContextUrls and the API reads the incoming header. That is the empty span in the checkout case above.
Vercel’s own dashboard shows traces under the Logs section through always-on tracing, which Vercel lists as Beta and available on all plans. Sending traces to your own backend through Trace Drains needs the Pro or Enterprise plan, and Trace Drains speak OTLP/HTTP only, not OTLP/gRPC. How spans leave a function before the instance stops is not stated in Vercel’s docs; the package picks the export mechanism for the environment.
The OpenTelemetry Collector: what it is, its configuration, Docker and the gateway pattern
The OpenTelemetry Collector is a separate service that receives, processes and exports telemetry, configured in one YAML file of receivers, processors, exporters and pipelines. It is optional: an SDK can export straight to a backend. A shared gateway Collector earns its place when credentials, sampling, redaction or a second destination need one home.
The project’s own advice leans the other way from mine, so here are both. The Collector docs say “in general we recommend using a collector alongside your service”, because the service can offload data quickly while the Collector handles “retries, batching, encryption or even sensitive data filtering”. On a managed host with one app, I skip it until there is a second destination or a redaction rule to enforce.
| Deployment | When I’d pick it | What you run |
|---|---|---|
| No Collector | One app, one backend, getting started | Nothing extra; the SDK exports straight to the backend, and the docs note that a change to collection, processing or ingestion requires code changes |
| Agent | You run your own hosts or containers | A Collector next to the app or on the same host, such as a sidecar |
| Gateway | Several services, shared credentials, sampling or redaction rules | One or more Collector instances as a standalone service behind a single OTLP endpoint |
The OTel gateway buys centrally managed credentials and one place for policy such as filtering or sampling. Its costs, in the docs’ words: “one more thing to maintain and that can fail”, added latency when Collectors are chained, and higher overall resource usage.
OpenTelemetry Collector configuration lives in a YAML file, by default /etc/<otel-directory>/config.yaml, and a component does nothing until a pipeline in the service section names it. Here is a minimal file with one OTLP receiver and one OTLP exporter, in the component names the configuration docs use:
receivers:
otlp:
protocols:
http:
endpoint: 0.0.0.0:4318
exporters:
otlp_http:
endpoint: https://otlp.example.com:4318
service:
pipelines:
traces:
receivers: [otlp]
exporters: [otlp_http]
The docs bind to 0.0.0.0 “as a convenience” and note the Collector defaults to localhost, which is safer when every client runs on the same host. Their examples batch inside the exporter with a sending_queue block containing batch, and the attributes processor can delete or hash a key such as email before export.
For an OTel Collector in Docker, the core distribution’s image is otel/opentelemetry-collector (also published on ghcr.io), and the contrib distribution’s is otel/opentelemetry-collector-contrib. The install docs mount your file over the default path with docker run -v $(pwd)/config.yaml:/etc/otelcol/config.yaml otel/opentelemetry-collector:0.162.0, version 0.162.0 being the one they printed on 3 October 2026.
Where traces go: Jaeger with OpenTelemetry, Jaeger vs Tempo, and an OpenTelemetry UI
OpenTelemetry leaves storage and viewing to other tools, so traces need a backend. Jaeger is an open-source tracing backend with its own UI that accepts OTLP. Grafana Tempo stores traces in object storage and is viewed through Grafana. Either is one more service to run, and hosted backends such as Datadog accept OTLP too.
OpenTelemetry has no visualization of its own, so the OpenTelemetry UI you end up using is whichever backend you send to. Jaeger with OpenTelemetry needs no Jaeger code in the app. Jaeger can receive data from the OpenTelemetry SDKs in their native OTLP, accepts only tracing data, and recommends the OpenTelemetry SDKs for instrumentation; its own clients “have been retired in 2022”. A Jaeger OTel setup is therefore the SDK, an exporter endpoint and a running Jaeger; with a Collector in between, the Collector’s OTLP exporter points at Jaeger instead of the SDK.
Jaeger’s getting-started docs run version 2.21 as one container, the all-in-one configuration with in-memory storage, with the UI at http://localhost:16686. That is for a first look on a laptop: Jaeger’s architecture page calls in-memory all-in-one “not recommended for production since the data is lost on restarts”, and the Badger storage option suits “only modest data volumes” on a single instance.
Grafana Tempo describes itself as a tracing backend that “only requires an object storage to operate”, read through Grafana’s built-in Tempo data source, and its repository states it is distributed under AGPL-3.0-only.
Jaeger vs Tempo, as I read the two docs: Jaeger is one tool with its own UI and search, Tempo is cheap storage that leans on Grafana for viewing, and each is a service you have to run, upgrade and back up. For Tempo vs Jaeger on a small team, the deciding question is usually whether Grafana is already running.
| Backend type | What you run | Storage | Rough effort for a small team |
|---|---|---|---|
| Jaeger all-in-one, in memory | One container with the collector and query roles | In memory, lost on restart | Low; for testing only |
| Jaeger with Badger | One instance | Badger, modest volumes | Moderate; no horizontal scaling |
| Jaeger with external storage | Collector and query roles plus a database such as Cassandra or Elasticsearch | The external database | High; two systems to run |
| Grafana Tempo | Tempo plus Grafana | Object storage | Moderate; Grafana must be running |
| Hosted backend | No server; an endpoint and a key | Held by the vendor | Low effort, metered bill |
What it costs to adopt, and when to skip it
Adopting OpenTelemetry has four costs: engineering time to instrument, a Collector or backend to run, ingest or storage fees that grow with traffic, and runtime overhead that sampling and batching keep down. I’d skip it while there is one service, no unexplained slow path and no error tracker yet.
| Cost line | What drives it | How to keep it small |
|---|---|---|
| Engineering time | Zero-code set-up is about an afternoon per service, by my estimate; useful custom spans are ongoing work | Start zero-code and add custom spans only where a question needs one |
| Something to run | A Collector and a self-hosted backend are two more services with their own upgrades and outages | No Collector and a hosted backend at first |
| Money | Hosted backends meter trace data: Datadog’s APM pricing counts ingested spans by the GB and indexed spans by the million, and Grafana Cloud meters traces by the GB | Sample, and switch off instrumentations that produce spans nobody reads |
| Runtime overhead | OpenTelemetry’s Java agent docs say a single overhead estimate is impossible and must be measured in your own system | Sampling and turning off unneeded instrumentations, both named in those docs |
Sampling is the lever on both money and overhead, and OpenTelemetry’s sampling page separates two kinds. Head sampling decides “as early as possible”, usually from the trace id and a target percentage, which is cheap but cannot promise to keep every trace with an error. Tail sampling decides after “considering all or most of the spans”, so it can keep every error trace and every slow one; it runs in the Collector’s Tail Sampling Processor and needs a stateful component that holds a large amount of data. The same page lists a case for not sampling at all: an app that generates “tens of small traces per second or lower”.
I’d add one more reason to wait: nobody will open the traces at least weekly. Until that changes, the money and hours go further on a request id written into every log line and on reading the database’s slow query view.
How to check your own app
An OpenTelemetry setup is verified with five checks: a known request found by its trace id, one tree of spans across every instrumented hop, the same id on that request’s log lines, no test email or prompt text in span attributes, and a forced error marked on its span.
Run them against a deployed staging or preview build, never the development server, with OTEL_TRACES_SAMPLER=always_on there so no test trace is sampled away. Two setups override that: on Vercel, Vercel’s own sampling rules must also agree before a span is emitted, and with Spring Boot’s own tracing the 10% default applies until you raise the probability.
- 01 Make one known request on staging and find its trace by id in the backend, allowing for the backend's own ingest delay. Fail: no trace, or the wrong service name. Keep: the trace URL.
- 02 Open the trace and confirm it is one tree: the inbound request, the database client if it is instrumented, and each outbound HTTP call appear as spans in the same trace. Fail: a second root span for the same request, which means the context was not passed at that hop. A hop with no span at all means that library is not instrumented. Keep: a screenshot of the test trace.
- 03 Find the same trace id on that request's log lines. This passes only when the logger writes the active trace id, for example as a trace_id field. Fail: log lines with no id or a different one. Keep: the log line beside the trace.
- 04 Sign in as a test user with a known test email, send one request carrying a known test prompt, then search the backend's span attributes for that exact email and that prompt text. Fail: any hit. Keep: the empty search result, dated.
- 05 Force an unhandled server error on a test route and confirm the span's status is Error with an event named exception. Fail: an Unset or Ok status on a failed request. Keep: the error span and your written sampling rule.
Check 3 works because OpenTelemetry SDKs can inject the trace id and span id into log records; if your logger does not, the trace and the logs never meet. Check 5 carries over to production only if the sampling rule keeps error traces: head sampling decides before the error happens, while tail sampling in a Collector can keep every trace that contains one. Write down which rule you run.
In the Production Hardening Sprint, deliverable 8.1, Structured, sanitized logs, is verified this way: “Trace a test request across services and check log content for sensitive fields.”
Where the sprint fits
No deliverable in the sprint is named OpenTelemetry. Deliverable 8.1 adds structured request logs with correlation IDs and appropriate user references, and excludes passwords, tokens, and unnecessary personal data. Another, 8.3, alerts on error spikes, latency, connection pressure, and queue backlog. A third, 8.8, creates one view of uptime, error rate, latency, signups, and revenue for the application’s relevant services. Hosting, paid tools, and API usage remain in your accounts. Each line is in the published scope.
Common questions about OpenTelemetry
Is OpenTelemetry difficult to learn?
No, not the concepts: span, trace, context propagation, SDK and Collector cover most of it. The difficulty is the number of moving parts, so start with zero-code instrumentation sending straight to a backend, with no Collector; OpenTelemetry’s Collector docs call that direct route a quick way to get value when getting started.
Is OpenTelemetry free or paid?
Free, and open source: the project’s repositories, including the specification, the Collector and the Python SDK, are published under the Apache 2.0 license. What costs money is the backend that stores the traces, if you use a hosted one, and the engineering time to instrument.
What is the difference between the OpenTelemetry Collector and an exporter?
An exporter is the component that sends telemetry to a destination; the Collector is a separate service that receives, processes and exports telemetry, using exporters of its own. Each SDK exports through an exporter, straight to a backend or to a Collector, and the Collector then forwards the data through the exporters its pipeline names, such as otlp_http in the configuration file above.
Is Jaeger deprecated?
No. Jaeger’s own client libraries were retired in 2022 in favor of the OpenTelemetry SDKs, but the Jaeger backend is maintained: its docs list version 2.21 as the latest release, dated 15 September 2026, and it accepts OTLP natively.
Is OpenTelemetry the same as Prometheus?
No. OpenTelemetry produces and ships telemetry of several kinds; Prometheus stores and queries metrics. The OpenTelemetry specification defines how Prometheus metric points translate to OTLP and back, so the two are built to work together.
Is OpenTelemetry the same as Grafana?
No. Grafana is where the data is viewed, and Grafana Labs also ships backends such as Tempo, which stores traces in object storage and shows them through Grafana’s built-in Tempo data source. OpenTelemetry produces the data those tools read.
What is the difference between an OpenTelemetry counter and a gauge?
A counter accumulates a value that only ever goes up, like a car’s odometer; a gauge measures a current value at the time it is read, like a fuel gauge. Between them sits the UpDownCounter, which accumulates but can also go down, such as a queue length.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase