HA Best Practices
Practice failover quarterly; apps must retry transient errors. These rules keep Patroni clusters actually available during real failures.
Search across all documentation pages
Practice failover quarterly; apps must retry transient errors. These rules keep Patroni clusters actually available during real failures.
mode: automatic). Reduces split-brain window.maximum_lag_on_failover from measured byte lag SLA. Block stale promotion./primary, not TCP 5432 alone. Prevents replica misroute.RECONNECT or reload PgBouncer in Patroni failover callback. Clear stale backends.server_lifetime low enough for failover (300s or less). Stale connection cap.connect_timeout in connection strings. Fail fast, retry sooner.patronictl switchover before OS patches on primary. Planned leader migration.pg_ctl promote outside Patroni. Avoid DCS state desync.patronictl pause; document owner. Resume automatic failover promptly.Patroni 3.x, 3-node etcd, HAProxy, PgBouncer, async physical replica, app retries, quarterly drill.
No for availability; yes if zero committed loss is an explicit RPO requirement.
RDS Multi-AZ trades control for ops simplicity. Patroni fits custom extensions and hooks.
Only promise what drills measure through the app connection string, plus buffer.
HA covers node/AZ loss. DR covers region loss. See DR Basics.
Stack versions: This page was written for PostgreSQL 18.4 (stable 18, maintenance 17), pgvector 0.8+, PgBouncer 1.x, Patroni 3.x, and PostGIS 3.5+.
Reviewed by Chris St. John·Last updated Jul 18, 2026