Diagnose Linux Memory Pressure and OOM Kills
When a Linux host runs low on memory, applications may slow down, allocations can fail, and the kernel may terminate a process through the out-of-memory (OOM) killer. A high memory reading alone does not prove a problem: Linux uses otherwise idle RAM for cache. Correlate symptoms with pressure, swap activity, kernel messages, and the affected process before changing limits or restarting services.
The commands in this guide are read-only. Run them during or soon after an incident when possible; process and pressure snapshots show only the current state, while kernel logs may preserve evidence of an earlier OOM event.
Step 1: Check Memory, Swap, and Pressure
Take a Read-Only Memory Snapshot
TriageCheck available memory and swap, then inspect the kernel’s pressure-stall information (PSI). PSI measures time tasks spend stalled for memory; sustained pressure is more useful than a single percentage showing RAM in use. vmstat reports memory and swap activity over one-second intervals after its initial average line.
free -hvmstat 1 5cat /proc/pressure/memory❯ View Expected Console Output
total used free shared buff/cache availableMem: 15Gi 13Gi 420Mi 1.1Gi 2.2Gi 1.4GiSwap: 4.0Gi 2.8Gi 1.2Gi
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu----- r b swpd free buff cache si so bi bo in cs us sy id wa st 2 0 2936012 430080 92160 2211840 12 64 120 210 800 1400 18 7 70 5 0
some avg10=4.21 avg60=2.87 avg300=1.30 total=123456789full avg10=1.14 avg60=0.82 avg300=0.35 total=34567890Step 2: Find Kernel OOM Evidence
Review Kernel Messages Around the Incident
Confirm OOMSearch the current boot’s kernel journal for OOM reports. If you know when the incident occurred, bound the query to a time window. The log may identify the killed process, its PID, memory usage, and the affected memory cgroup. Lack of a match does not rule out an OOM event if logs rotated, the host rebooted, or the kernel ring buffer was lost.
sudo journalctl -k -b --no-pager | grep -Ei 'out of memory|oom-kill|killed process'
# Narrow the search to an incident window (adjust the timestamps).sudo journalctl -k -b --since '2026-10-06 09:00:00' --until '2026-10-06 09:30:00' --no-pager❯ View Expected Console Output
kernel: memory: usage 2048MB, limit 2048MB, failcnt 37kernel: Memory cgroup out of memory: Killed process 2481 (worker) total-vm:3120040kB, anon-rss:1842200kBStep 3: Identify Current Memory Consumers
Compare Process Memory with System Totals
Locate UsageSort processes by resident memory to find large current consumers. RSS is a useful lead, but it can count shared pages in more than one process and will not explain all kernel or cgroup memory. Capture the owning service and workload before deciding whether a process is expected to use that amount.
ps -eo pid,ppid,user,comm,rss,%mem --sort=-rss | head -n 15❯ View Expected Console Output
PID PPID USER COMMAND RSS %MEM 2481 2100 app worker 1842200 11.4 902 1 root java 1324000 8.2Step 4: Check the Service’s Cgroup Limit
Compare Service Usage with Its Limit
Check ScopeA host can have available RAM while a service is OOM-killed because its cgroup has a lower memory limit. For a systemd service, inspect its control group and resource properties. On cgroup v2 systems, read memory.current, memory.max, and memory.events from that group’s directory. Replace the example unit with the service involved; permissions may require sudo.
systemctl show example.service -p ControlGroup -p MemoryCurrent -p MemoryMax
# Use the ControlGroup path printed above beneath /sys/fs/cgroup.sudo cat /sys/fs/cgroup/system.slice/example.service/memory.currentsudo cat /sys/fs/cgroup/system.slice/example.service/memory.maxsudo cat /sys/fs/cgroup/system.slice/example.service/memory.events❯ View Expected Console Output
ControlGroup=/system.slice/example.serviceMemoryCurrent=2147483648MemoryMax=2147483648
low 0high 0max 42oom 3oom_kill 2Step 5: Preserve Evidence and Choose a Response
Record the incident time, affected service, kernel OOM lines, free and vmstat output, PSI values, and cgroup counters. Compare with the service’s normal workload and recent changes. If the cgroup counters show OOM kills while host memory remained available, investigate whether the configured limit matches the workload. If host-wide pressure is sustained, identify the workload and memory trend before considering capacity, application tuning, or workload scheduling changes.
Avoid killing large processes or raising memory limits as an automatic response: either action can interrupt work or shift pressure to the whole host. Validate any proposed limit change against host capacity, competing services, and the application’s documented requirements, then monitor PSI, swap activity, and OOM counters after the change.