An OOM kill is the Linux kernel ending a process because the machine or its container ran out of memory. The process gets signal 9, which it cannot catch, so there is no stack trace and nothing in your error tracker: only exit code 137 and one kernel log line naming the process and its memory figures.

OOM kill: what it is, and the one log line it leaves

An OOM kill is the Linux kernel stopping one process with signal 9 because memory ran out on the machine or inside the container’s limit. The process cannot catch the signal, so it exits with code 137 and leaves no stack trace. The kernel log names the process it chose.

Memory is one of the ceilings every piece of web performance optimization eventually meets, and running out of it ends the process instead of slowing it down. The kernel’s own account in the kernel’s memory management concepts is short: on a loaded machine the kernel can “be unable to reclaim enough memory to continue to operate”, so it invokes the OOM killer, which “selects a task to sacrifice for the sake of the overall system health”. A process that was OOM killed received SIGKILL, and the signal(7) manual page lists SIGKILL among the signals that “cannot be caught, blocked, or ignored”. So the process gets no moment to log, flush a buffer or report to anything, which is why the error tracker inside it stays quiet: its code never runs again. The bash manual gives the exit code: a command “terminated by signal n” returns 128+n, and SIGKILL is signal 9, so the status is 137.

The Linux kernel writes the OOM kill record to its log, and in its own spelling the event is oom-kill, as the second line below shows. The format comes from the kernel source (mm/oom_kill.c) of the 7.3 development tree in September 2026, with every value replaced by a placeholder:

Out of memory: Killed process <pid> (<name>) total-vm:<n>kB, anon-rss:<n>kB, file-rss:<n>kB, shmem-rss:<n>kB, UID:<uid> pgtables:<n>kB oom_score_adj:<adj>
oom-kill:constraint=<constraint>,nodemask=...,task=<name>,pid=<pid>,uid=<uid>

Read it field by field. <pid> and <name> are the victim. anon-rss is its resident anonymous memory, the kind the kernel docs say is created “for program’s stack and heap”, so for a Node or Python app it is mostly the heap; file-rss and shmem-rss are resident file mappings and shared memory; total-vm is the total program size, much of which the process may never have touched. oom_score_adj is the adjustment covered under the score below. A kill inside a container limit starts with Memory cgroup out of memory: instead of Out of memory:, and its summary line reads constraint=CONSTRAINT_MEMCG. Kernels 5.4, 6.6 and 6.12 print the same kill line; 4.19 prints a shorter one with no Out of memory: prefix, no UID and no oom_score_adj. What the line never tells you is why the memory grew.

What is OOM? OOM meaning in software

OOM in software means out of memory: a program, or its machine, asked for more RAM than it could have. It shows up three ways: an error the program can catch, such as Java’s OutOfMemoryError; a kernel kill it cannot catch; and a GPU memory error in model code. Only the second is an OOM kill.

The same three letters cover all three, so whether an OOM is an error you can catch or a kill you cannot depends on who reported it:

  • From the runtime: Java throws java.lang.OutOfMemoryError, and Python raises MemoryError “when an operation runs out of memory but the situation may still be rescued”. Node sits in between: when its heap reaches its cap it aborts with “FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory”, which your code cannot catch but which at least prints why.
  • From the kernel: the OOM kill this page is about: no exception, one log line, exit code 137.
  • From the GPU framework: PyTorch raises an out of memory error when a CUDA allocation cannot be met, a different problem from your server’s RAM.

Outside software OOM stands for other things entirely; the one a founder is likely to meet is covered in the questions at the end.

Why it matters for a small app: a crash with no stack trace, shown five ways

The same kill looks different on each platform, and each one describes it in its own words. What your users see is my reading and the same everywhere: requests in flight fail, the app is gone while it restarts, then it comes back and the cycle repeats as memory climbs again.

Where the app runsWhat the owner sees, in the platform’s wordsWhere to lookWhat the user sees
Plain Linux server or VPSOut of memory: Killed process <pid> (<name>) in the kernel logjournalctl -k or dmesgFailed requests, then nothing until a supervisor starts the process again
DockerOOMKilled true in the container state, “since the container was last started”, beside the last exit code; an oom eventdocker inspect, docker eventsFailed requests while the container restarts, if a restart policy is set
Kubernetes”Last State: Terminated”, “Reason: OOMKilled”, “Exit Code: 137”, and a rising “Restart Count”kubectl describe podFailed requests on that pod while Kubernetes restarts the container
RenderThe exact out-of-memory event wording is not stated in Render’s docs; they say “warnings about resource constraints usually appear in the service’s logs and on the service’s Events page”Events page; Metrics page, memory as a percentage of the planFailed requests during the restart
Heroku”Error R14 (Memory quota exceeded)”: the dyno pages to swap and keeps running. “Error R15 (Memory quota vastly exceeded)”: the dyno “will be forcibly killed with SIGKILL”App logs; the Metrics tab’s memory chartHeroku’s “standard error page with the HTTP status code 503”

Render’s wording comes from Render’s troubleshooting guide, which also suggests scaling the service when resources run short; its metrics page shows memory “as a percentage of the maximum allowed value for your service’s compute plan”. Heroku’s two codes are in Heroku’s error codes: only R15 ends in a kill; R14 isn’t currently emitted for Fir-generation apps, and R15 isn’t planned for them. What those failed requests look like to a visitor, the gateway and unavailable error pages, is the ground covered in what a crash looks like when an app goes viral.

The reason this goes unnoticed is ordinary. 17 of the 21 third-party apps had no error tracking or alerting: when a user hits an error, nothing records it. Those 21 are the third-party apps I audited in June and July 2026, a selected set of audited apps, not a random sample and not a rate for AI-built apps in general. Even an app with an error tracker gets nothing here, because a process killed with SIGKILL has no chance to send a report, so the signal has to come from outside the process; that split is why uptime monitoring and error tracking do different jobs.

One look-alike is worth ruling out first. An app that stops answering under load while its memory graph stays flat is usually out of database connections, not memory (my reading), and connection pool exhausted errors and their fixes covers that failure. Which ceiling an app reaches first, memory or connections or something else, is the question behind why an AI app stalls at 100 concurrent users.

How it works: the kernel’s choice, the score, the container limit, and the fix order

Four parts, in the order you need them: what sets the killer off, how it picks a process, why a container changes the picture, and what to change.

The Linux OOM killer: what triggers it

The Linux OOM killer runs when the kernel cannot free enough memory for a request it already promised. Linux overcommits memory on purpose, reclaims page cache and swaps first, and kills a process only when reclaim fails. Swap delays that moment and slows the machine. It does not remove the cause.

The OOM killer is the kernel’s last resort, and what triggers it is failed reclaim, not memory pressure on its own. Linux hands out more memory than it has, a feature the kernel’s vm sysctl documentation says can be very useful because “there are a lot of programs” that allocate huge amounts “just-in-case” and “don’t use much of it”. As free memory drops, the kernel frees page cache and swaps out anonymous memory; at a lower threshold, the min watermark, an allocation “will trigger direct reclaim” and is “stalled until enough memory pages are reclaimed to satisfy the request”. Only when that fails does the kill happen. On Linux the log also shows which process asked for the memory, as <name> invoked oom-killer, and that process need not be the one that dies. Swap gives reclaim somewhere to put memory, so the kill comes later, but a process that keeps growing keeps growing, now on a slower machine; Docker’s own docs call swap “slower than memory”.

Forum answers about OOM in Linux often reach for three settings. They set machine-wide policy, and none of them fixes a single app’s memory:

  • vm.overcommit_memory: 0, the default, “rejects obvious overcommits”; 1 “pretends there is always enough memory until it actually runs out”; 2 is a “never overcommit” policy.
  • vm.panic_on_oom: 1 makes the kernel panic instead of killing; the docs say values 1 and 2 “are for failover of clustering”.
  • vm.oom_kill_allocating_task: non-zero kills “the task that triggered the out-of-memory condition” instead of scanning for the best victim.

Some systems also run a userspace killer. The systemd-oomd manual describes a service that watches cgroups and pressure stall information to “take corrective action before an OOM occurs in the kernel space”, sending SIGKILL to every process in the cgroup it picks, so on a machine that runs it the kill may come from there and not from the kernel.

The OOM score: how the kernel picks your process

The OOM score is the number in /proc/<pid>/oom_score, and the process with the highest score is killed first. It mostly tracks the share of allowed memory a process uses, shifted by oom_score_adj, which runs from -1000 (never kill) to 1000. On a one-app server the app is the biggest process, so it is the usual victim.

The kernel’s proc filesystem docs describe the heuristic as a value “ranging from 0 (never kill) to 1000 (always kill)”, whose units “are roughly a proportion along that range of allowed memory the process may allocate”. Allowed memory is the configured limit when a memory limit was reached, and “all allocatable resources” when the whole system ran out. The file you read already includes the adjustment, “so it is effectively in range [0,2000]”. Read both files for any process:

cat /proc/<pid>/oom_score
cat /proc/<pid>/oom_score_adj

Because the app dominates memory on a single-app box, tuning scores rarely helps a small SaaS, and I’d keep it for one case only. That case is a box where the database or the SSH daemon runs beside the app: lowering their scores means the app is killed instead of them, and you keep a way in. The value -1000 “is equivalent to disabling oom killing entirely for that task”. Setting it on the app itself does not create memory, so the kernel kills something else, possibly the database. Docker’s docs warn against the same move, telling you not to set --oom-score-adj “to an extreme negative number on the daemon or a container”.

Containers and managed hosts: when the limit is the cgroup, not the machine

A container OOM kill happens at the cgroup’s memory limit, not the machine’s: the kernel kills inside that group even when the host has free memory. Docker records OOMKilled beside exit code 137, Kubernetes reports the reason OOMKilled, and a managed host’s plan RAM is the same kind of limit.

A container’s limit is a cgroup setting. In the cgroup v2 documentation, memory.max is the “memory usage hard limit”: “If a cgroup’s memory usage reaches this limit and can’t be reduced, the OOM killer is invoked in the cgroup.” The group keeps count in memory.events: oom counts the times usage “reached the limit and allocation was about to fail”, and oom_kill counts “processes belonging to this cgroup killed by any kind of OOM killer”. Docker’s resource constraints page sets the limit with -m or --memory and warns about what can happen without one: when the host runs short, “Any process is subject to killing, including Docker and other important applications.” Kubernetes’ guide to container resources sets it with resources.limits.memory and says limits are enforced reactively: a container over its limit “may not be immediately killed”, and when the killed process is the container’s PID 1 and the container is restartable, “Kubernetes restarts the container.” On a managed host, the RAM of the plan you pay for plays the same role, which is why Render shows memory as a share of the plan.

Kill typeWhat ran outWhere the limit is setHow to tell
Machine-level killThe server’s RAM, after reclaim and swapThe machine itself; nothing per processKill line starts Out of memory: (check 1)
Container limit killThe cgroup’s memory.maxdocker run --memory, or resources.limits.memory in KubernetesMemory cgroup out of memory: in the kernel log, oom_kill in memory.events, OOMKilled in the state (checks 1 and 2)
Managed-host limit killThe memory of the plan or instance typeThe plan you pick on Render or the dyno type on HerokuThe host’s own message, such as Heroku’s R15, and a memory graph at the plan’s maximum (check 1’s dashboard look)

The trap for garbage-collected runtimes is a heap sized from the wrong number. If the runtime sizes its default heap from the machine’s memory, the heap can grow past the container’s limit and the kernel kills the process before the runtime ever throws its own out-of-memory error. The JDK 21 java command reference says the JVM handles this: its container support “is enabled by default” and lets it “determine the amount of memory and number of processors that are available to a Java process running in docker containers”, and its default -XX:MaxRAMPercentage is “25 percent”. Node’s docs say only that the default heap limit is “determined by system resources”; whether a container limit counts as one is not stated in Node’s docs, so check 3 below prints the real figure instead of trusting it. What happens after the kill, a restart loop or a container marked unhealthy, comes down to Docker Compose health checks and restart policies. A serverless function has the same limit under another name, its configured memory size, which sits with Lambda concurrency and the limits that bite first.

How to avoid the OOM killer: what to change, in order

The way to avoid the OOM killer is six changes in order: see whether memory climbs or spikes, fix the spike in the code, fix the climb in the code, cap the runtime’s heap below the container limit, cap concurrency, and only then buy more memory. Disabling the killer or adding swap hides the cause.

The order is my working rule: the cheapest and most certain changes come first.

  1. 01 Find the shape. Read the platform's memory graph over a day: a line that climbs and never comes back down is a leak or an unbounded cache; a flat line with tall spikes is one heavy request. Confirm it by matching the spikes to requests in your logs.
  2. 02 Fix the spike at its source: stream files and exports instead of reading them whole, cap the page size a caller may ask for, process uploads and AI responses in chunks, and move heavy jobs off the request path. Confirm it by replaying the heavy request and watching the spike shrink.
  3. 03 Fix the climb: bound every in-process cache and map, remove listeners and timers that are never cleared, and compare two heap snapshots taken some time apart to see what grew. Confirm it when the graph's floor after each collection stops rising.
  4. 04 Cap the runtime below the container limit, so the runtime fails first with its own error instead of a silent kill: --max-old-space-size in Node, -XX:MaxRAMPercentage on the JVM, worker recycling with max_requests under Gunicorn. Confirm it with check 3 below.
  5. 05 Cap concurrency: fewer workers, or a limit on requests in flight per instance. Confirm it by running the load test again and reading peak memory at the same traffic.
  6. 06 Only then buy memory: move to a bigger instance or plan. Confirm it by watching a full day of the memory graph at the new size.

One real finding shows the cap in step 2 going missing. In one of my own apps, a messages route accepted any page size with no cap. The audit could not say how likely abuse was, because the route had no API documentation and no other caller in the repository; it recorded that the severity rises if the API is documented or opened to outside integrators. The lesson I take from it: a page size the caller picks is, in my reading, a memory size the caller picks, so the cap belongs on the server. Moving the heavy job out of the request is how to run long tasks in the background; choosing which rows and columns a page reads is the oversized-reads section of the article on apps that stall under load, above, and it is not repeated here.

For step 3, Node can write a heap snapshot on a signal: Node’s command-line options describe --heapsnapshot-signal, which “causes the Node.js process to write a heap dump when the specified signal is received”, so two dumps some time apart show what grew.

For step 4, --max-old-space-size takes a size “in MiB” and “Sets the max memory size of V8’s old memory section”; Node’s docs give the example of setting it to 1536 “On a machine with 2 GiB of memory” to leave room for other uses. My working rule for a one-process Node app is a cap of roughly 70 to 80 percent of the container’s limit, lower when the process holds large buffers outside the heap. On the JVM the matching setting is -XX:MaxRAMPercentage, with the default of 25 percent the java command reference gives. Python has no heap-size option like Node’s: its command-line documentation lists none. The usual guard there is worker recycling, and Gunicorn’s settings describe max_requests as “a simple method to help limit the damage of memory leaks”, with max_requests_jitter there “to stagger worker restarts to avoid all workers restarting at the same time”. Both default to 0, which disables the restarts. The same logic covers how to fix an OOM error that comes from the runtime itself, such as Node’s heap abort: that error is the cap doing its job, and steps 2 and 3 are the real fix.

For step 5, memory use grows with the requests in flight, so fewer at once lowers the peak (my reading). For step 6, a bigger instance is a fair fix for an app that has simply grown; the choice between one bigger instance and several smaller ones is horizontal scaling vs vertical scaling, and the way the same pressure shows up as slow requests before any kill is tail latencies and p99.

Three things do not belong on the list: disabling the killer, setting the app’s oom_score_adj to -1000, and adding swap and calling it fixed.

How to check your own app

An OOM kill fix is proven with five checks: the last kill line found and read, the container’s OOMKilled state read beside exit code 137, a heap cap below the container limit, a staging load test at production’s limit that ends with no kill, and memory and restart alerts that have each fired once.

Start in the browser. On Render, open the service’s Events page and its Metrics page, where memory can be shown as a share of the plan; on Heroku, open the Metrics tab, where “Memory quota is depicted as a dashed gray line with any quota breaches flagged in red” (application metrics are not available on eco dynos). A restart with a memory reason beside a sawtooth graph that touches the ceiling is a strong lead before you open a terminal.

  1. 01 Find the last kill. On a server you run, search the kernel log with journalctl -k or dmesg for the kill line and note the process, the time and the memory figure; journalctl -k reads only the current boot unless you pass another, and dmesg may need extra rights where the kernel restricts it. On a managed host, the events page or app log is the source. Pass: you have the line. Fail: no line or event near the restart's time, and check 2 decides. Evidence: the line and its time.
  2. 02 Read the state. In Docker, inspect the container's OOMKilled flag and exit code, and read docker events for an oom event too. In Kubernetes, describe the pod and read the last state's reason and exit code. Pass for memory as the cause: OOMKilled true, an oom event, or reason OOMKilled. Exit code 137 alone does not prove memory. Fail: none of the three, so memory is not shown as the cause. Evidence: the command output.
  3. 03 Print the real heap limit inside the running container, with the same NODE_OPTIONS production uses. Pass: a limit below the container's memory limit with the headroom of your working rule. Fail: a limit at or above it, or a cap set in a file the process never reads. For the JVM, compare what it reports against the limit; for Python under Gunicorn, pass is max_requests set and workers times peak memory per worker, measured in check 4, below the limit. Evidence: the printed figure and the limit.
  4. 04 Reproduce in staging, never in production: run the load test against staging with the same memory limit as production, once before the change and once after. The run before must show the kill or the climb; if it cannot, raise the load or lengthen the run until it does. Pass after the change: a run at the expected peak with no kill line, no OOMKilled, and a memory floor that stops rising after each collection. Fail: a kill line, OOMKilled, or a floor that keeps rising at the same load. Evidence: both reports with workload, duration, concurrency, latency, error rate and peak memory.
  5. 05 Confirm someone would know: set a memory alert and a restart alert, lower each threshold in staging below today's value so it fires once, then restore it. Where the host offers no such alert, an outside uptime monitor's down alert catches a restart that lasts long enough to fail its checks. Pass: two alert messages. Fail: a threshold crossed with no message. Evidence: the messages with their times.

The commands for checks 1 to 3, all read-only:

journalctl -k | grep -i 'killed process'
dmesg -T | grep -i 'killed process'
docker inspect --format '{{.State.OOMKilled}} {{.State.ExitCode}}' <container>
docker events --since 24h --filter event=oom
docker exec <container> node -e "console.log(require('v8').getHeapStatistics().heap_size_limit / 1048576)"

A few notes on reading them. dmesg falls under the kernel’s dmesg_restrict switch: when it is 1, “users must have CAP_SYSLOG to use dmesg(8)”, and its -T option prints readable times that its manual warns “could be inaccurate”. As the Docker row of the table says, the flag covers only the current run, so a container in a restart loop can show false after a kill; docker events keeps the oom event, though “Only the last 256 log events are returned”. Exit code 137 alone proves only SIGKILL, because docker stop sends SIGTERM and, “after a grace period, SIGKILL” (Docker’s docker stop reference). The last line prints heap_size_limit, which Node’s v8 module defines as “the maximum size of the V8 heap, in bytes”, divided here into MiB; docker exec runs with the environment variables “set at the time the container is created”, so it sees the container’s NODE_OPTIONS. For the JVM, the java command reference lists -XshowSettings:vm, which “Shows the settings of the JVM”, and -XshowSettings:system, which shows “host system or container configuration”.

Two of the checks carry my working rules. For check 4, I’d run staging about as long as the usual gap between kills in production, or an hour when kills come faster, and a run that cannot show the problem before the change proves nothing after it. For check 5, I’d set the memory alert at about 80 to 90 percent of the limit. A threshold set just above today’s value never fires, which is why each alert is fired on purpose. Hosts differ on what they offer: Render’s notification list includes “A running service becomes unhealthy” but has no memory event, and Heroku’s threshold alerting sets limits on 95th percentile response time and the percentage of failed requests, only on Professional and Fir dynos, while its Metrics tab shows online notifications whose examples “include alerts on memory errors”. Check 4 is load testing a web application at production’s memory limit, check 5 rests on knowing how to alert on error rate spikes and other thresholds, and the wider checklist these checks sit inside is scaling web applications.

Keep the dated set together: the kill line, the state output, the heap figure, the before and after reports, and the alert messages.

Where the sprint fits

In the Production Hardening Sprint, deliverable 9.2, load test and bottleneck fixes, means we simulate concurrent users, identify the first bottlenecks, fix them, and rerun the workload, and we verify it this way: report the workload, duration, environment, concurrency, latency, and error rate before and after changes. Deliverable 6.6, background processing, moves long-running AI generation, exports, and bulk communication into background jobs, and we verify it by running a long task beyond the normal request window and verifying completion and user-visible status. The fee covers the engineering work; hosting, paid tools, and API usage remain in your accounts, and any required third-party costs are explained before they are enabled. Both deliverables are listed with the rest in the published scope.

Common questions about running out of memory

What is kill 9 and kill 15?

Kill 15 sends SIGTERM, a request to stop that a process can handle; kill 9 sends SIGKILL, which “cannot be caught, blocked, or ignored”. The OOM killer uses SIGKILL, which is why the process leaves no trace of its own. docker stop sends 15 first and 9 after the grace period, so a container stopped slowly also exits with 137 without any memory problem.

How to fix OOMKilled in Kubernetes?

Fix OOMKilled in Kubernetes by finding out why the container outgrew its limit before you raise the limit. Run kubectl describe pod and read the last state, compare real memory use with the limit, fix the app’s use first, set the runtime’s heap cap below the limit, and only then set the request and the limit from measured usage (that order is my working rule). Kubernetes’ own docs say the next step might be to check the application code for a memory leak, and if it behaves as expected, “consider setting a higher memory limit (and possibly request) for that container”.

What is an OOM in AI?

In model code, an OOM in AI usually means the GPU ran out of memory, which PyTorch reports as an out of memory error. PyTorch’s CUDA notes call one allocator setting “particularly useful for inference serving, where a fatal GPU OOM would crash the server process”: with it, the serving framework can “catch the OutOfMemoryError and reject the individual request while continuing to serve subsequent requests”. An app that calls a hosted model API holds no GPU memory of its own, though buffering large responses whole can still exhaust the app’s own RAM.

What does OOM stand for in business?

In business, OOM stands for order of magnitude, a rough sense of size, and has nothing to do with memory. Wikipedia’s OOM page lists order of magnitude beside out of memory, and The Sugar Engineers’ cost estimation guide uses “Order-Of-Magnitude (OOM) Estimate” for an estimate “generally used by client or management in feasibility studies”.

What is OOM risk?

OOM risk is the chance that a workload goes past its memory limit and gets killed. In my reading it rises with no limit set, no heap cap, unbounded in-process caches, and per-request memory that grows with the size of the input, such as a page size or an upload the caller picks.