The Linux CLI
A PostgreSQL server does not run in isolation, it runs as a process on a Linux host that has its own disk, memory, and network limits, so reading the command line is really reading a second layer underneath the database itself.
Search across all documentation pages
A PostgreSQL server does not run in isolation, it runs as a process on a Linux host that has its own disk, memory, and network limits, so reading the command line is really reading a second layer underneath the database itself.
This page builds the mental model for that reading: the three layers a triage session correlates, the doctrine that keeps an operator from changing things blindly, and where the CLI's evidence stops and a database-level decision begins.
Every triage session on a Postgres host is really working across three layers at once, even when it doesn't feel that way.
The first layer is the operating system: disk space, memory pressure, and network connections, all visible through tools like df, free, and ss that know nothing about SQL.
The second layer is the process: the actual postgres workers running on that host, visible through ps, pgrep, and systemctl, which know a process exists but not what query it's running.
The third layer is database state: pg_stat_activity, pg_locks, and the server logs, which know exactly what SQL is running but nothing about the host's free memory.
A useful analogy is a doctor reading vital signs, an X-ray, and a patient interview separately - each layer is accurate on its own terms, and a diagnosis only holds once all three agree.
The core discipline this page teaches is correlation: a disk-full symptom at the OS layer and a PANIC: could not write to file in the Postgres log are the same event seen from two layers, not two separate problems.
Learning to move fluently between these three layers, rather than living in just one of them, is what separates fast triage from guesswork.
The reason CLI triage has a strict order - OS first, then process, then database - is that each layer's tools answer a different time window and a different kind of question.
OS-level commands like vmstat and iostat answer "what is happening right now on this host," independent of which database or query caused it.
Process-level commands like ps and systemctl status answer "is the Postgres process itself healthy," which matters because a restart loop can masquerade as a query performance problem.
Database-level views like pg_stat_activity answer "what is Postgres doing right now," including wait events that explain why a backend is slow rather than just that it's slow.
The triage doctrine that ties these together is simple: measure at all three layers before changing any configuration, restarting any service, or killing any process, because acting on one layer's evidence alone routinely misdiagnoses the other two.
This is also why an incident bundle script exists - capturing df, free, vmstat, and the relevant pg_stat_activity/pg_locks snapshots together, in one pass, before anything about the system changes.
# Minimal three-layer snapshot, taken together, not sequentially over minutes
df -h "$PGDATA"; free -h
ps -eo pid,rss,cmd --sort=-rss | grep postgres | head -5
psql "$DATABASE_URL" -c "SELECT wait_event_type, count(*) FROM pg_stat_activity GROUP BY 1;"Captured together, these three commands tell a coherent story; captured five minutes apart during a live incident, they can each describe a different moment.
The doctrine of "measure before you touch" becomes harder to hold to under real incident pressure, which is exactly when it matters most.
A common failure mode is an operator seeing high CPU or IO and reaching straight for a GUC change or a failover, when the three-layer read would have shown a transient checkpoint spike that was already resolving on its own.
Escalation to failover is a database-level decision that CLI evidence can support but should rarely trigger alone - hardware failure, corruption, or an unrecoverable disk are the kinds of OS-layer findings that actually justify it.
The distinction between a graceful stop and a kill is where CLI mechanics meet real risk: systemctl stop lets Postgres complete its checkpoint, while kill -9 on the postmaster risks an incomplete checkpoint and a longer, riskier recovery on next start.
Access control follows the same layering logic - read-only OS and SQL commands are reasonable for any engineer to run, while anything that stops a process or changes a GUC belongs behind an on-call or buddy-system gate.
As infrastructure shifts toward containers, the same three-layer model still applies inside a kubectl exec session, it just replaces systemctl with the container runtime's own process visibility.
| Approach | Strength | Weakness | Best Fit |
|---|---|---|---|
Single-layer read (only pg_stat_activity, say) | Fast, requires no host access | Misses host-level root causes entirely, like disk or memory exhaustion | Quick sanity check when you already trust host health |
| Full three-layer correlation | Catches root causes a single layer would miss | Takes longer and requires broader access | Any real incident, or any symptom you can't explain from one layer alone |
| Act first, measure after | Feels fast under pressure | Frequently misdiagnoses transient spikes as real problems and causes unnecessary failovers | Never - included here only as the failure mode this page argues against |
kill -9 is a fast, safe way to restart a stuck Postgres" - it risks an incomplete checkpoint, turning a slow restart into a longer, riskier recovery.pg_stat_activity alone tells you if the host is healthy" - it only knows about SQL-level state, and says nothing about disk space, memory pressure, or the process's own restart history.df, free, vmstat, and ps are present on nearly every Linux image and cover most first-pass triage without any extra install.The operating system (disk, memory, network), the Postgres process itself (running, restarting, or crashed), and the database's own internal state (queries, locks, wait events).
Because system state changes minute to minute during an incident, and commands run five minutes apart can each be describing a different moment rather than one coherent story.
A graceful stop lets Postgres complete its current checkpoint before shutting down, while kill -9 interrupts that process and risks a longer, riskier recovery on the next start.
When it points to something a restart can't fix, like hardware failure, filesystem corruption, or an unrecoverable disk - not for an ordinary lock storm or transient IO spike.
Read-only OS and SQL commands are reasonable for most engineers, while anything that stops a process or changes configuration should require an on-call role or a buddy.
Yes - kubectl exec gets you into the same three layers, it just substitutes the container runtime's process visibility for systemctl.
Checkpoint writes, autovacuum, and host-level contention can all produce a similar CPU signature, which is exactly why single-layer conclusions are unreliable.
df, free, vmstat, and ps cover most first-pass triage and are present on nearly every Linux image without extra installation.
The pressure to skip it is exactly when the risk of misdiagnosis is highest, which is why the doctrine exists as a default rather than a suggestion.
At the OS layer it's df reporting 100% usage, at the process layer it may show as a restart loop, and at the database layer it surfaces as a PANIC in the server log.
Acting - restarting, killing a process, or failing over - on evidence from only one layer, before correlating it against the other two.
Stack versions: This page was written for PostgreSQL 18.4 running on Linux (systemd, Debian and RHEL layouts); it is otherwise conceptual and not tied to a specific tool version.
Reviewed by Chris St. John·Last updated Jul 15, 2026