Primary unresponsive or no Leader: recover it without losing data
Use this when the cluster has no Leader, or the primary is unreachable over SSH or its metrics
scrapes are failing. Do not fail over first.
A failover promotes a replica without the writes it never received. If the old primary’s Postgres is still running (for example, because only its OS disk or Patroni is stuck), the old primary is the most complete copy of the data. This runbook recovers it in place, so no committed writes are lost.
Timebox: assessment ≤ 10 min.
Before you start: is the watchdog active?
Section titled “Before you start: is the watchdog active?”On any cluster member (the configuration is the same across the cluster):
ls -l /dev/watchdog; lsmod | grep softdog; sudo grep -A3 '^watchdog:' /var/opt/gitlab/patroni/patroni.yml- Not active (no device, no module, no
watchdog:block): follow this runbook. Nothing will fence the primary automatically. You have to. - Active: a frozen primary is reset by its own kernel once Patroni stops checking in for the
Patroni TTL minus
safety_margin(about a minute and a half with our current settings), and a replica is promoted within a couple of minutes, before you can pause. Expect a bounded loss of the writes only the old primary had. Go to Step 5, and keep the old primary’s data disk.
Step 1: Disable chef, then pause Patroni
Section titled “Step 1: Disable chef, then pause Patroni”First disable chef on the healthy members, so a chef run can’t revert the pause. Exclude the
stuck primary: knife ssh would hang on it.
knife ssh 'roles:<cluster-role> AND NOT name:<primary-fqdn>' 'sudo chef-client-disable "INC-xxxx: paused for recovery"'Then pause Patroni, from any healthy member. Pausing stops any replica from promoting. It never touches Postgres.
sudo gitlab-patronictl pause --waitsudo gitlab-patronictl list | grep 'Maintenance mode'While paused there is no automatic failover for any reason, and the watchdog is inactive. Keep the pause as short as the assessment and recovery.
Step 2: Assess (5 minutes)
Section titled “Step 2: Assess (5 minutes)”Grafana queries use datasource mimir-gitlab-gprd and are written for the main cluster. For
other clusters, change type (patroni-ci, patroni-sec, patroni-registry) and, in the
PgBouncer query, fqdn to that cluster’s PgBouncer hosts. For gstg, use env="gstg" and
mimir-gitlab-gstg.
Open the checks together (all in Explore, main, last 6h), or one at a time from the links in the table.
| # | Question | Check | How to read it |
|---|---|---|---|
| 0 | Has Consul lost the primary? | From any node: consul members | grep -E 'failed|left' and consul kv get service/<scope>/leader (<scope> is scope: in patroni.yml) | A failed primary plus a missing leader key means the lock has expired. That’s why there’s no Leader. It’s usually the first signal. |
| 1 | Are the replicas receiving WAL? | Grafana: count(pg_stat_wal_receiver_status{env="gprd", type="patroni"}) or vector(0) vs count(pg_replication_is_replica{env="gprd", type="patroni"} == 1). On a replica: select status, sender_host, flushed_lsn, last_msg_receipt_time from pg_stat_wal_receiver; | A detached replica’s series disappears, so the first count drops below the second (or to 0). No row on a replica means it’s detached. This is when writes start existing only on the primary. |
| 2 | Is the old primary still taking writes? | Direct (preferred), from any replica, twice, 10 s apart: sudo gitlab-psql -h <primary-fqdn> -Atc "select now(), pg_current_wal_lsn()". SSH to the primary isn’t needed. Fallback (incomplete): Grafana sum(rate(pgbouncer_stats_sql_transactions_pooled_total{env="gprd", fqdn=~"pgbouncer-0[0-9]-db-gprd.*", database="gitlabhq_production"}[2m])) covers web and API only. | If the LSN moves, the primary is still writing. A low PgBouncer rate doesn’t mean writes have stopped. If you can’t check, assume writes are continuing. |
| 3 | How far behind are the replicas? | Grafana: max by (fqdn) (time() - pg_stat_wal_receiver_last_msg_receipt_time{env="gprd", type="patroni"}). On each replica: select pg_last_wal_receive_lsn(), pg_last_xact_replay_timestamp(); | How much exists only on the primary. Don’t compute byte lag from Prometheus, because scrape timing makes it unreliable. |
| 4 | Is WAL reaching GCS? | Grafana: max by (fqdn) (pg_archiver_pending_wal_count{env="gprd", type="patroni"}) | Normally < 20; thousands means archiving is stuck, so no replica can catch up from the archive. |
| 5 | Is the primary’s Patroni half-dead? | curl -s --max-time 3 http://<primary>:8009/patroni | jq '{state, role, lock_age_s: ((now|floor) - .dcs_last_seen)}'. The TTL is in sudo gitlab-patronictl show-config | grep ttl | role master/primary with lock_age_s below the TTL: healthy. Above the TTL: half-dead. It’s writing, has lost the lock, and nothing will fence it. A timeout means Patroni is dead or the node is frozen; rely on check 2. |
Record on the incident: when the replicas detached, whether the primary is still writing, and how far behind the replicas are.
Step 3: Recover the old primary in place
Section titled “Step 3: Recover the old primary in place”-
Save evidence if you can. A reset loses what hasn’t been written to disk:
gcloud compute instances get-serial-port-output <vm> --project=<project> --zone=<zone> > <vm>-serial.log -
Reset the VM with
gcloud compute instances reset <vm> --project=<project> --zone=<zone>or from the GCP console, or restart Postgres if the node is still responsive. Committed transactions are safe in the WAL on the data disk. While paused, Patroni removes the leader lock once Postgres stops, but no replica promotes. -
As soon as the node is reachable, disable chef on it, so a scheduled chef run doesn’t touch Patroni:
sudo chef-client-disable "INC-xxxx: paused for recovery" -
Start Postgres by hand. Patroni doesn’t start it while paused:
Terminal window sudo -u gitlab-psql /usr/lib/postgresql/<ver>/bin/pg_ctl -D /var/opt/gitlab/postgresql/data<ver> startpg_ctl: another server might be runningis expected after a crash (a stalepostmaster.pid). -
Confirm it’s Leader on the same timeline:
Terminal window grep -E "PAUSE: acquired session lock|PAUSE: no action" /var/log/gitlab/patroni/patroni.log | tail -3sudo gitlab-patronictl list # the old primary = Leader, TL unchangedThe replicas may start streaming from it again before you resume. That’s expected.
-
Only then resume, and re-enable chef on all members:
Terminal window sudo gitlab-patronictl resume --waitknife ssh 'roles:<cluster-role>' 'sudo chef-client-enable'
⚠️ Never resume while there is no Leader. A replica will promote immediately, and every write that only the old primary has will be lost.
Step 4: If the node doesn’t come back
Section titled “Step 4: If the node doesn’t come back”Stay paused. While the cluster is paused, no replica promotes and the data stays on the old primary’s disk. Nothing is lost while you troubleshoot.
- Use the GCP serial console and the VM’s logs to find out why it isn’t booting or why Postgres won’t start.
- Bring in a second DBRE to help. Don’t delete or detach the old primary’s disks.
A failover is out of scope for this runbook. It loses every write that only the old primary has, so it’s a data-loss decision: don’t make it alone. It needs a second DBRE to review and agree, with the Incident Manager informed.
Step 5: After recovery
Section titled “Step 5: After recovery”- Delayed and archive replicas: if the old primary’s WAL past a fork reaches GCS, these replicas
can follow the old timeline and then stop with
new timeline N forked off current database system timeline N-1 before current recovery point. They then need a rebuild. Check their position (select pg_last_wal_replay_lsn();) against the fork LSN in the new timeline’s.historyfile. Restarting a replica before it reaches the fork, once that.historyfile is in GCS, should make it follow the new timeline. That last step hasn’t been tested yet, so confirm with a DBRE first. See also Diverged timeline WAL in GCS. - Check that application traffic has recovered: error rates, and that the PgBouncer pools point at the primary.
Notes and limitations
Section titled “Notes and limitations”- Assumes asynchronous replication. With synchronous replication, a failover normally loses no confirmed writes, so recovering in place matters less. Update this runbook if that changes.
- While paused, nothing is automatic. No failover happens for any failure, and the watchdog is inactive. Keep the pause short, and have someone watching the cluster the whole time.
- A VM reset is an unclean shutdown. Postgres runs crash recovery on start, which can take minutes on a busy primary. Transactions that hadn’t committed are rolled back, and client connections drop. Committed transactions are not affected.
- Replicas catch up from the primary’s replication slots. If a slot is missing, or the WAL it
needed was removed (
max_slot_wal_keep_size), that replica needs a rebuild. The primary’s data is unaffected. - Testing so far: validated on a 2-node test cluster. Crash-recovery time and catch-up time at production scale haven’t been measured yet.
Validation
Section titled “Validation”The recover-in-place procedure has been tested in the db-benchmarking environment: a VM reset of the primary under write load, with replication cut. No rows were lost and there was no timeline change. A game day in staging is planned.