Skip to content

Primary unresponsive or no Leader: recover it without losing data

Use this when the cluster has no Leader, or the primary is unreachable over SSH or its metrics scrapes are failing. Do not fail over first.

A failover promotes a replica without the writes it never received. If the old primary’s Postgres is still running (for example, because only its OS disk or Patroni is stuck), the old primary is the most complete copy of the data. This runbook recovers it in place, so no committed writes are lost.

Timebox: assessment ≤ 10 min.

On any cluster member (the configuration is the same across the cluster):

Terminal window
ls -l /dev/watchdog; lsmod | grep softdog; sudo grep -A3 '^watchdog:' /var/opt/gitlab/patroni/patroni.yml
  • Not active (no device, no module, no watchdog: block): follow this runbook. Nothing will fence the primary automatically. You have to.
  • Active: a frozen primary is reset by its own kernel once Patroni stops checking in for the Patroni TTL minus safety_margin (about a minute and a half with our current settings), and a replica is promoted within a couple of minutes, before you can pause. Expect a bounded loss of the writes only the old primary had. Go to Step 5, and keep the old primary’s data disk.

First disable chef on the healthy members, so a chef run can’t revert the pause. Exclude the stuck primary: knife ssh would hang on it.

Terminal window
knife ssh 'roles:<cluster-role> AND NOT name:<primary-fqdn>' 'sudo chef-client-disable "INC-xxxx: paused for recovery"'

Then pause Patroni, from any healthy member. Pausing stops any replica from promoting. It never touches Postgres.

Terminal window
sudo gitlab-patronictl pause --wait
sudo gitlab-patronictl list | grep 'Maintenance mode'

While paused there is no automatic failover for any reason, and the watchdog is inactive. Keep the pause as short as the assessment and recovery.

Grafana queries use datasource mimir-gitlab-gprd and are written for the main cluster. For other clusters, change type (patroni-ci, patroni-sec, patroni-registry) and, in the PgBouncer query, fqdn to that cluster’s PgBouncer hosts. For gstg, use env="gstg" and mimir-gitlab-gstg.

Open the checks together (all in Explore, main, last 6h), or one at a time from the links in the table.

#QuestionCheckHow to read it
0Has Consul lost the primary?From any node: consul members | grep -E 'failed|left' and consul kv get service/<scope>/leader (<scope> is scope: in patroni.yml)A failed primary plus a missing leader key means the lock has expired. That’s why there’s no Leader. It’s usually the first signal.
1Are the replicas receiving WAL?Grafana: count(pg_stat_wal_receiver_status{env="gprd", type="patroni"}) or vector(0) vs count(pg_replication_is_replica{env="gprd", type="patroni"} == 1). On a replica: select status, sender_host, flushed_lsn, last_msg_receipt_time from pg_stat_wal_receiver;A detached replica’s series disappears, so the first count drops below the second (or to 0). No row on a replica means it’s detached. This is when writes start existing only on the primary.
2Is the old primary still taking writes?Direct (preferred), from any replica, twice, 10 s apart: sudo gitlab-psql -h <primary-fqdn> -Atc "select now(), pg_current_wal_lsn()". SSH to the primary isn’t needed. Fallback (incomplete): Grafana sum(rate(pgbouncer_stats_sql_transactions_pooled_total{env="gprd", fqdn=~"pgbouncer-0[0-9]-db-gprd.*", database="gitlabhq_production"}[2m])) covers web and API only.If the LSN moves, the primary is still writing. A low PgBouncer rate doesn’t mean writes have stopped. If you can’t check, assume writes are continuing.
3How far behind are the replicas?Grafana: max by (fqdn) (time() - pg_stat_wal_receiver_last_msg_receipt_time{env="gprd", type="patroni"}). On each replica: select pg_last_wal_receive_lsn(), pg_last_xact_replay_timestamp();How much exists only on the primary. Don’t compute byte lag from Prometheus, because scrape timing makes it unreliable.
4Is WAL reaching GCS?Grafana: max by (fqdn) (pg_archiver_pending_wal_count{env="gprd", type="patroni"})Normally < 20; thousands means archiving is stuck, so no replica can catch up from the archive.
5Is the primary’s Patroni half-dead?curl -s --max-time 3 http://<primary>:8009/patroni | jq '{state, role, lock_age_s: ((now|floor) - .dcs_last_seen)}'. The TTL is in sudo gitlab-patronictl show-config | grep ttlrole master/primary with lock_age_s below the TTL: healthy. Above the TTL: half-dead. It’s writing, has lost the lock, and nothing will fence it. A timeout means Patroni is dead or the node is frozen; rely on check 2.

Record on the incident: when the replicas detached, whether the primary is still writing, and how far behind the replicas are.

  1. Save evidence if you can. A reset loses what hasn’t been written to disk: gcloud compute instances get-serial-port-output <vm> --project=<project> --zone=<zone> > <vm>-serial.log

  2. Reset the VM with gcloud compute instances reset <vm> --project=<project> --zone=<zone> or from the GCP console, or restart Postgres if the node is still responsive. Committed transactions are safe in the WAL on the data disk. While paused, Patroni removes the leader lock once Postgres stops, but no replica promotes.

  3. As soon as the node is reachable, disable chef on it, so a scheduled chef run doesn’t touch Patroni: sudo chef-client-disable "INC-xxxx: paused for recovery"

  4. Start Postgres by hand. Patroni doesn’t start it while paused:

    Terminal window
    sudo -u gitlab-psql /usr/lib/postgresql/<ver>/bin/pg_ctl -D /var/opt/gitlab/postgresql/data<ver> start

    pg_ctl: another server might be running is expected after a crash (a stale postmaster.pid).

  5. Confirm it’s Leader on the same timeline:

    Terminal window
    grep -E "PAUSE: acquired session lock|PAUSE: no action" /var/log/gitlab/patroni/patroni.log | tail -3
    sudo gitlab-patronictl list # the old primary = Leader, TL unchanged

    The replicas may start streaming from it again before you resume. That’s expected.

  6. Only then resume, and re-enable chef on all members:

    Terminal window
    sudo gitlab-patronictl resume --wait
    knife ssh 'roles:<cluster-role>' 'sudo chef-client-enable'

⚠️ Never resume while there is no Leader. A replica will promote immediately, and every write that only the old primary has will be lost.

Stay paused. While the cluster is paused, no replica promotes and the data stays on the old primary’s disk. Nothing is lost while you troubleshoot.

  • Use the GCP serial console and the VM’s logs to find out why it isn’t booting or why Postgres won’t start.
  • Bring in a second DBRE to help. Don’t delete or detach the old primary’s disks.

A failover is out of scope for this runbook. It loses every write that only the old primary has, so it’s a data-loss decision: don’t make it alone. It needs a second DBRE to review and agree, with the Incident Manager informed.

  • Delayed and archive replicas: if the old primary’s WAL past a fork reaches GCS, these replicas can follow the old timeline and then stop with new timeline N forked off current database system timeline N-1 before current recovery point. They then need a rebuild. Check their position (select pg_last_wal_replay_lsn();) against the fork LSN in the new timeline’s .history file. Restarting a replica before it reaches the fork, once that .history file is in GCS, should make it follow the new timeline. That last step hasn’t been tested yet, so confirm with a DBRE first. See also Diverged timeline WAL in GCS.
  • Check that application traffic has recovered: error rates, and that the PgBouncer pools point at the primary.
  • Assumes asynchronous replication. With synchronous replication, a failover normally loses no confirmed writes, so recovering in place matters less. Update this runbook if that changes.
  • While paused, nothing is automatic. No failover happens for any failure, and the watchdog is inactive. Keep the pause short, and have someone watching the cluster the whole time.
  • A VM reset is an unclean shutdown. Postgres runs crash recovery on start, which can take minutes on a busy primary. Transactions that hadn’t committed are rolled back, and client connections drop. Committed transactions are not affected.
  • Replicas catch up from the primary’s replication slots. If a slot is missing, or the WAL it needed was removed (max_slot_wal_keep_size), that replica needs a rebuild. The primary’s data is unaffected.
  • Testing so far: validated on a 2-node test cluster. Crash-recovery time and catch-up time at production scale haven’t been measured yet.

The recover-in-place procedure has been tested in the db-benchmarking environment: a VM reset of the primary under write load, with replication cut. No rows were lost and there was no timeline change. A game day in staging is planned.