FilesystemExt4Errors
Overview
Section titled “Overview”The kernel has recorded ext4 filesystem errors on a Gitaly node. ext4 keeps a
persistent error counter in the filesystem superblock and increments it when it
detects metadata corruption, a failed journal commit, or an I/O error it cannot
transparently retry. The counter survives remounts and reboots, so a non-zero
reading is a record of damage rather than a transient condition.
On a Gitaly node the affected mount usually holds git repository data, so corruption can surface to users as failed clones, fetches or pushes against the repositories on that node, and in the worst case as data loss. The node may also have remounted read-only, which fails every write to it.
The alert is a cause alert: it reports a hardware or filesystem fault, not degraded user-facing SLIs. It is expected to be rare — over the 30 days to 2026-09-11 it did not fire at all, and no node’s counter increased.
The recipient is expected to identify the affected node and mount, decide whether the node is still safe to serve traffic from, and drain it if it is not. Do not treat a firing as noise: nothing increments this counter under normal operation.
Services
Section titled “Services”- Service Overview
- Team that owns the service: Tenant Scale:Gitaly Team
Metrics
Section titled “Metrics”The alert reads filesystem_ext4_errors_total, a node-exporter counter of the
ext4 superblock error count, labelled by fqdn, device and mountpoint. The
unit is a count of errors, not a rate.
increase(filesystem_ext4_errors_total{env="gprd", service="gitaly"}[5m]) > 0- The threshold is any increase, held for 5 minutes. There is no tuned
value to choose: the counter only moves when the kernel has recorded a real
error, so any movement is the signal. The 5-minute
forexists to survive a single bad scrape, not to filter benign activity. - Under normal conditions every series reads a constant value — usually 0 — and
increase()is 0. As of 2026-09-11 there were 492 series across 164 Gitaly nodes ingprd, andmax(increase(...[30d]))was 0. - A node that has been reimaged can start a new series at 0. A node that previously recorded errors keeps its non-zero total; the alert fires on the increase, so a historical non-zero value does not re-fire by itself.
Alert Behavior
Section titled “Alert Behavior”- Expected to be rare.
ALERTS{alertname="FilesystemExt4Errors"}had no samples over the 30 days to 2026-09-11. - Silence only after the affected node has been drained or replaced, and scope
the silence to that
fqdn. A blanket silence on the alertname hides genuine corruption on every other node.
Severities
Section titled “Severities”- Configured as s3, so it does not page. It reports damage to one node, not a service-wide outage.
- Impact is limited to the projects whose repositories live on the affected node, so it is a subset of customers — but for that subset the impact can be severe (failed git operations, and potentially unrecoverable data).
- Raise severity if the node has remounted read-only, if more than one node is affected at once, or if Gitaly error-rate or apdex SLIs for the shard are also degraded. A single node with a small counter bump and healthy SLIs stays s3.
Verification
Section titled “Verification”-
The query that triggered the alert:
increase(filesystem_ext4_errors_total{env="gprd", service="gitaly"}[5m]) -
Take the affected node and mount from the
fqdn,deviceandmountpointlabels on the firing series. -
On the node, the kernel’s own record is authoritative and more detailed than the counter:
Terminal window dmesg -T | grep -i -E 'ext4|i/o error|remount'sudo dumpe2fs -h /dev/<device> | grep -i -E 'error|state|mount'mount | grep -w /var/opt/gitlab # is it ro?dumpe2fs -hreportsFS Error count, the first and most recent error, and the filesystem state. A state other thancleanmeans the filesystem needs attention before the node serves traffic again.
Troubleshooting
Section titled “Troubleshooting”Cheapest first:
- Confirm the mount is still read-write. A read-only remount is the case that needs immediate action, because every write to that node is already failing.
- Read
dmesgfor the underlying cause. An I/O error points at the disk or the GCE persistent disk / local NVMe beneath it; a metadata error with no I/O error points at filesystem corruption. - Check whether the counter is still climbing, or recorded a single event and stopped. A one-off with a clean filesystem state and healthy SLIs is a different problem from a node actively accumulating errors.
- Check the Gitaly SLIs for the affected shard, to establish whether users are being served errors from this node right now.
- Check whether other nodes in the same zone or of the same machine type are affected, which would suggest an infrastructure-wide fault rather than one bad disk.
Do not run fsck on a mounted filesystem, and do not resize or repartition
the volume to “clear” the counter. The counter is a record; clearing it destroys
the evidence without fixing the cause.
Possible Resolutions
Section titled “Possible Resolutions”- Single transient I/O error, filesystem clean, SLIs healthy. Record it and monitor. No action is a valid outcome here.
- Counter still climbing, or filesystem not clean. Drain the node and move its repositories off, then have the disk replaced. See move-repositories and new-storage.
- Node remounted read-only. Treat as an outage for the repositories on that node: drain it immediately rather than waiting for a repair.
- Repository-level corruption reported by users on this node. See gitaly-repository-corruption.
Escalation
Section titled “Escalation”Gitaly has a Tier 2 rotation. Follow the How to Escalate guidance to page the team.
Several Slack channels are also available: