Which logs severity levels should stay switched on in production? Info, warn, error and fatal stay on. Debug and trace stay off until a question needs them; then, by my rule, they come on for one service, for about an hour, through a setting. Syslog uses its own scale, 0 to 7, with 0 the most severe.
Logs severity levels: the ladder from trace to fatal
Log severity levels are labels that rank each log line by how much it matters. The common ladder, from least to most severe, is trace, debug, info, warn, error and fatal. Setting a logger to info writes info and everything above it, and ignores debug and trace.
A level is a label on each line that tells the log store how much the line matters, so the store can filter, keep, drop or alert on lines without reading them. Read the log levels in that order, trace at the bottom and fatal at the top: setting one rung writes that rung and everything more severe. Python’s reference states the rule in its own words: messages “which are less severe than level will be ignored”. Log4j 2 lists the same six names as its standard levels, FATAL, ERROR, WARN, INFO, DEBUG and TRACE, plus OFF and ALL as switches rather than labels. What each level is for, judged by the response a line needs and with worked examples, comes down to choosing log levels by the response they need. This page is the production policy that sits on top of that, one part of the logging and monitoring a small SaaS needs.
The table adds what a policy needs and a definition does not: whether the level is on in production, and who hears about it.
| Level | The question it answers | On in production by default | Who is told |
|---|---|---|---|
| trace | Which line ran, step by step? | Off | Nobody |
| debug | What did the code see: inputs, branch taken, query built? | Off | Nobody |
| info | What happened: a sign-up, a job, a webhook accepted? | On | Nobody; it is read when needed |
| warn | What nearly went wrong and recovered? | On | Someone, in a weekly look |
| error | What failed for a user? | On | An alert |
| fatal | Why did the process stop? | On | A page to whoever is on call |
The last two columns reflect my practice, not a standard.
Trace sits below debug and follows execution line by line; OpenTelemetry’s log data model describes TRACE as “A fine-grained debugging event. Typically disabled in default configurations.” Setting a loglevel of trace, where the logger has one, writes every log line the code contains, which is why it is the first level to leave off.
Not every logger has all six. Python names five levels, has no trace, and calls its top level CRITICAL. Go’s log/slog names four. The numbers also run in opposite directions: in Python, Go and OpenTelemetry a bigger number is more severe, while Log4j 2 gives FATAL 100 and TRACE 600, and syslog puts its most severe level at 0. So compare levels by name, never by number.
Syslog severity levels 0 to 7, and what a syslog facility is
Syslog severity levels are 8 numbers listed in RFC 5424, most severe first: 0 Emergency, 1 Alert, 2 Critical, 3 Error, 4 Warning, 5 Notice, 6 Informational and 7 Debug. The syslog facility is a second number, 0 to 23, naming which part of the system wrote the message.
Syslog is the message protocol that Linux servers, network gear and many log drains speak, and RFC 5424, titled “The Syslog Protocol”, is the document that describes it. Every syslog message carries one of these severity levels at the start of its header. The RFC prints the two tables “for purely informational purposes”, calling the values “not normative but often used”.
Below are the syslog severity codes with the RFC’s own descriptions, the short keyword journalctl accepts for each, and the app logger level each one lands on in OpenTelemetry’s published example mapping, which is the simplest way to read syslog log levels beside the ones your code writes.
| Code | Keyword | RFC 5424 description | Nearest app logger level (OpenTelemetry mapping) |
|---|---|---|---|
| 0 | emerg | Emergency: system is unusable | fatal |
| 1 | alert | Alert: action must be taken immediately | error |
| 2 | crit | Critical: critical conditions | error |
| 3 | err | Error: error conditions | error |
| 4 | warning | Warning: warning conditions | warn |
| 5 | notice | Notice: normal but significant condition | info |
| 6 | info | Informational: informational messages | info |
| 7 | debug | Debug: debug-level messages | debug |
Syslog has no trace level. OpenTelemetry maps only Emergency to FATAL and places Alert and Critical higher inside its ERROR range, so in that mapping a syslog Critical counts as a severe error, not a fatal one. OpenTelemetry’s log data model defines its own SeverityNumber from 1 to 24, where “Smaller numerical values correspond to less severe events”, and it is the neutral reference wherever two scales meet.
The facility answers a different question: where the message came from. RFC 5424 lists codes 0 to 23, such as 0 for kernel messages, 2 for the mail system and 4 for security/authorization messages, and leaves 16 to 23, local0 to local7, for local use. There are no syslog facility levels as such: the facility number names the source, and the severity number beside it carries the level.
The two numbers travel together as the priority value in angle brackets at the start of each message: the facility times 8, plus the severity. The RFC’s own example is a “local use 4” message (facility 20) at Notice (severity 5), which gives 20 times 8 plus 5, written as <165>.
A web app owner meets this scale on a Linux server, a managed database or a log drain, where the store filters on these numbers. On a systemd server, a journal entry can carry a PRIORITY= field from 0 (“emerg”) to 7 (“debug”), “compatible with syslog’s priority concept”, and journalctl filters on it. The rest of the syslog message format belongs with structured logging, not with levels.
What debug logs are, and what debug logging means
Debug logs are lines written for the developer who will diagnose a problem later: the inputs a function received, the branch it took, the query it built, the response it got. Debug logging means running with the level set to debug. Most loggers ship with a higher default, so debug lines are not written until someone lowers it.
The purpose is to answer a question after the fact, without the code running in front of you. Python describes its DEBUG level as “Detailed information, typically only of interest to a developer trying to diagnose a problem”, and its root logger is created at WARNING, so debug lines stay silent until the level comes down. Node apps usually get their levels from a library, as with pino logging in a Node app, and Python apps from the standard module, as with the Flask logger and Python log levels.
A useful debug line carries the request id, named fields rather than a sentence with values glued in, and no secrets. A debug line earns its place only if you can find it again: search by request id, and every line that request wrote is there, in order.
Log debugging is reading those lines instead of attaching a debugger. In production it is usually the only option, because nobody pauses a live server to step through a request.
When an app’s settings ask you to enable debug logging
A debug logging switch in a consumer app’s settings tells the app to record a detailed diagnostic log on your device so its support team can read it. Discord, for example, puts the switch under Voice & Video, Debugging, and uploads the log to its support team when you ask it to.
So turning on “enable debug logging” in an app you use means the app writes a fuller debug log file for its own developers, not that you are changing anything on a server. Discord’s bug-report article gives the desktop path: open User Settings with the cogwheel, select Voice & Video and the Debugging tab, toggle Debug Logging on, reproduce the bug, then select Upload under Debug Logging; uploaded logs “are sent automatically to Support”. The same article names the folders where the files sit: C:\Users\*username*\AppData\Roaming\discord\logs on Windows and ~/Library/Application Support/discord/ on macOS.
My advice for that file: treat it as private, send it only through the vendor’s own upload, and switch the setting off once support has it. The rest of this page is for the people who run the app that writes the logs.
Why it matters: what debug in production costs a small SaaS
Debug left on in production costs money and can expose data, and levels set wrong hide real failures. The table puts each cost next to the level rule that prevents it.
| What goes wrong | How it shows up | The level rule that prevents it |
|---|---|---|
| Volume | The log bill grows with every request | Production runs at info; debug comes on for one service with an end time |
| Exposure | Request bodies, headers and query parameters land in the log store, sometimes with passwords and tokens inside | Debug lines are sanitized before debug is ever switched on |
| Noise | Everything logged as error mutes the error alert, and handled failures logged at info trigger nothing | Error means a failure a user felt; a handled failure that still matters goes to warn, never info |
Volume is the cost you can price. Amazon CloudWatch charges $0.50 per GB of standard log data ingested in US East (N. Virginia), per Amazon CloudWatch pricing as read on October 3, 2026. Take an app with 1,000,000 requests a day, 2 lines per request at info, 20 at debug, 300 bytes a line, a 30-day month and 1 GB as a billion bytes: info writes 18 GB a month, about $9 of ingestion, and debug writes 180 GB, about $90, before storage and before the free tier. Those figures are my assumptions for the arithmetic, and a debug line that prints a payload can run far past 300 bytes, which only raises the debug total.
Exposure is the cost you cannot undo, because debug is where code prints what it received. Keeping secrets out of every line is part of how to do logging with one event shape, and the debug level is where that rule gets tested first.
Noise matters only if something is listening. Among the apps I audited, 17 of the 21 third-party apps had no error tracking or alerting: when a user hits an error, nothing records it. I audited those 21 apps in June and July 2026 and chose them myself, so the count is not a rate for AI-built apps in general. A level can only help an alert that exists.
Picture a founder chasing a webhook bug who adds a debug line printing each incoming payload, lowers the production level to debug through an environment variable, finds the bug and ships the fix. The level stays at debug, so every request keeps writing its debug lines, the log store keeps billing for every gigabyte ingested, and the payloads, with whatever personal data the provider put in them, stay in the store for as long as it keeps logs. The switch back was nobody’s job.
How it works: debug versus info, setting the level, and switching debug on
Three questions come up in order: which level a line belongs at, how the level is set in your runtime, and when debug may come on.
Log debug vs info: where the line sits
Log debug vs info comes down, for me, to one test: info is a fact about the business or the system that someone will want next month, and debug is a fact about the code that someone wants only while chasing a bug. Both calls stay in the code; the configured level decides which lines are written.
Eight events from a typical SaaS, sorted by that test:
| Event | Level | Why |
|---|---|---|
| A password reset email was queued | info | Support will ask about it next month |
| An invoice PDF was generated | info | A business fact with a customer behind it |
| A webhook event was accepted, with its event id | info | The id ties the line to the provider’s own record |
| The SQL text a query ran | debug | A fact about the code, useful only while chasing one bug |
| A feature flag was evaluated for one user | debug | Explains one branch for one request |
| The raw response body from an AI provider | debug | Useful in a bug hunt, and a leak risk: sanitize or leave it out |
| A provider rate-limited the call and a retry was scheduled | warn | Nearly went wrong, recovered |
| A user typed a wrong password | info, or nothing | The user caused it, so it is never an error |
logger.info() is the call that writes a line at info, the level Python’s docs describe as “Confirmation that things are working as expected.” logger.debug() and logger.info() cost about the same to call; whether the line gets written depends on the configured level. With the log level at debug, both kinds of line are written; logging at info writes only the info lines. So debug calls stay in the code, and the setting controls them.
One catch: a debug line that builds an expensive value still pays for it at info. Python’s logger takes the arguments separately, as in logger.debug("query %s", sql), and merges them into the message with msg % args only when a handler formats the record; for work that is costly to produce, guard it with logger.isEnabledFor(logging.DEBUG).
Setting the level: go log level, java.util.logging and the rest
Log level names differ by runtime: Go’s slog has 4 levels stored as integers, java.util.logging has 7 from SEVERE down to FINEST, and Python names 5. In each, my rule is to read the level from one environment variable, and in Java remember that the handler’s level must allow the line too.
| Runtime | Level names, least to most severe | Default | How to set it |
|---|---|---|---|
Go, log/slog | DEBUG (-4), INFO (0), WARN (4), ERROR (8) | Info | HandlerOptions.Level, or a LevelVar to change it while running |
Go, log | No levels in its documented index: Print, Fatal and Panic functions | Not applicable | Not applicable |
Java, java.util.logging | FINEST, FINER, FINE, CONFIG, INFO, WARNING, SEVERE (plus OFF and ALL) | ConsoleHandler level defaults to INFO | logger.setLevel(Level.FINE), or a .level line in logging.properties, plus the handler’s level |
| Java, Log4j 2 | TRACE (600), DEBUG (500), INFO (400), WARN (300), ERROR (200), FATAL (100) | Not stated on Log4j’s levels page | The Log4j configuration |
Python, logging | DEBUG (10), INFO (20), WARNING (30), ERROR (40), CRITICAL (50) | Root logger created at WARNING | logger.setLevel(...) |
| Node.js | Set by the logging library you add | The library’s default | The library’s option, read from an environment variable |
In Go’s log/slog package the levels are integers with gaps, the default is Info, and a LevelVar “allows the level to be varied dynamically”, while the older log package documents no levels at all. This reads the level from one variable and keeps a handle to change it later:
level := new(slog.LevelVar) // Info by default
if v := os.Getenv("LOG_LEVEL"); v != "" {
if err := level.UnmarshalText([]byte(v)); err != nil {
level.Set(slog.LevelInfo)
}
}
h := slog.NewJSONHandler(os.Stderr, &slog.HandlerOptions{Level: level})
slog.SetDefault(slog.New(h))
// at run time: level.Set(slog.LevelDebug)
UnmarshalText accepts the level names in any case, so LOG_LEVEL=debug works.
In java.util.logging the seven levels run from SEVERE at the top to FINEST at the bottom, and “Enabling logging at a given level also enables logging at all higher levels.” When you set the level on a Java logger to FINE and still see no FINE lines, look at the handler: a handler discards “Message levels lower than this value”, and a ConsoleHandler defaults to INFO. Log4j 2’s level list gives each name a number, and apart from OFF at 0 the smallest belongs to FATAL; choosing between Java logging frameworks is a bigger topic than levels. AWS Lambda’s own level controls belong with AWS logging and monitoring.
Python has the same trap: Python’s logging reference says a logger’s lines go out “unless a handler’s level has been set to a higher severity level”.
On every runtime, I read the level from one environment variable and never hard-code it.
Should debug logging be on or off?
Debug logging should be off in production and on in development. I’d switch it on in production only as a deliberate act: for one service, through a setting rather than a code change, with an end time of about an hour and a named owner, after confirming the debug lines carry no secrets.
These are the five rules I work to:
- 01 Production runs at info.
- 02 The level comes from an environment variable or a runtime flag, so raising it needs no code change. On hosts where a changed environment variable only applies after a redeploy or restart, count that step in the time.
- 03 Debug goes on for one service or one module, where the logger supports it, not for the whole app.
- 04 It goes on with an end time, about an hour, and a named person switches it back.
- 05 Before it goes on, confirm the debug lines are sanitized, because this is the moment request bodies reach the log store.
A finer tool than a whole service is debug for one request: a header or user flag that raises the level for that request alone, where the logger’s child loggers allow it. You then find that request afterwards by a trace id on every line.
The levels left on have their own jobs. Error and fatal lines feed alerting, which is a question of how to alert on error rate spikes. How long each level is kept is a retention choice: how long to keep application logs. In a container, debug volume is also a disk question, which comes down to how to clear Docker logs before they fill the disk.
How to check your own app
Log levels are checked with 5 tests: no debug lines in production over a day the log store still holds, a line count by level that looks plausible, debug switched on and off through the setting, a forced failure that lands at error, and a search of the debug lines for the test values you sent.
Run checks 1 and 2 in production and checks 3 to 5 in staging set to production’s level. Each one can fail, and each leaves evidence to keep.
- 01 In the production log store, count info lines over the last 24 hours, then count debug lines over the same window. Expected: some info, zero debug. If the info query returns nothing, the store does not hold that window (retention shorter than a day, or logs not shipped) and the zero debug count proves nothing. Keep both counts and the window.
- 02 Count one day's lines by level. An error share above a few percent, or zero errors on an app with real traffic, is a prompt to look at how lines are labeled, not a standard. Keep the counts.
- 03 Raise the level to debug through the setting, then sign in once with a test account whose email and password you chose. Find that request's debug lines by its request id. Set the level back, sign in again, and confirm no new debug lines appear. Note how long each change took to apply, and keep both searches with timestamps.
- 04 Point one provider's base URL at an address that refuses connections, send one request that calls that provider, let any retries run out, and confirm the failure lands at error with the request id, not at info. Keep the line.
- 05 Search the debug lines from check 3 for the exact test email, the test password and the session token that sign-in returned. Expected: zero hits for the password and the token. Keep the search and its count.
What each search looks like depends on your log store; the counts and the request id are what matter. Sprint deliverable 8.1, Structured, sanitized logs, is verified this way: “Trace a test request across services and check log content for sensitive fields.”
Where the sprint fits
The Production Hardening Sprint covers this page’s ground in three deliverables. Deliverable 8.1 is structured, sanitized logs: we add structured request logs with correlation IDs and appropriate user references, and exclude passwords, tokens, and unnecessary personal data. Under 8.3, operational threshold alerts, we alert on error spikes, latency, connection pressure, and queue backlog. With 8.5, log retention, we set and document log retention across the application’s services. The fee covers the engineering work; hosting, paid tools, and API usage remain in the client’s accounts, and any required third-party costs are explained before they are enabled. Each deliverable and its verification step is listed in the published scope.
Common questions about log levels
How do I stop debug logs?
Set the level back to info in the one place it is read, usually an environment variable or a runtime flag, restart or redeploy if your host needs that for the change to apply, then search for debug lines written after the change and confirm there are none. In a consumer app such as Discord, turn the Debug Logging switch off in the same settings screen.
How to check debug logs?
In your log store, filter by level debug and by the request id of the request you are chasing. On a Linux server with systemd, journalctl -p debug..debug -u <service> shows only debug entries for one service; a single -p debug shows every level from debug up, because one value means “this log level or a lower (hence more important) log level”. In a consumer app, open the folder the vendor’s own article names, such as Discord’s logs folder.
Where are debug logs stored?
Debug logs go wherever every other level goes: the standard output your host captures, the log store it ships to, or on a Linux server the systemd journal, kept on disk under /var/log/journal or in memory under /run/log/journal depending on its Storage= setting in the journald.conf manual. journald’s MaxLevelStore= defaults to “debug”, so it stores debug entries unless someone lowers it. A consumer app writes its debug log to a file on your device.
What is log level 4?
In syslog, level 4 is Warning, “warning conditions”. In Go’s slog, 4 is the integer behind WARN. Other loggers use other numbers for the same idea (Python’s WARNING is 30), so configure levels by name.
What is local 0 to 7 in syslog facility?
local0 to local7 are syslog facility codes 16 to 23, which RFC 5424 reserves for local use, so your own programs can be told apart from the kernel, mail or auth messages on the same server. A local4 message at Notice gets the priority value 165.
What are the downsides of using syslog?
Syslog does not confirm delivery, and over an unreliable transport such as UDP “some messages may be lost”. Receivers must accept messages of up to 480 octets and should accept 2048; a message longer than a receiver supports may be truncated or discarded. Messages are free text unless the sender adds structured data, and the scale has no trace level. For a web app, JSON lines to standard output are simpler to search and to ship.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase