Host: kolu-ci-3 (kolu CI venue pool, x86_64-linux, Incus container)
Status: still live and taking CI work — pull it from the pool now
Verdict: one unclean shutdown two weeks ago, not a failing disk. Recycle the container.
Correction: the first version of this gist said "one specific store path is damaged." That was wrong about scope, and acting on it would not have fixed anything — see below. The symptom and the host are right; the blast radius was understated by three orders of magnitude.
ci::home-manager fails in ~1.2 s:
error: error parsing derivation '/nix/store/0mvnbryv5wvkryyw2m67xhnx5vj359gk-initrd-linux-6.18.39.drv': file is empty (possible filesystem corruption)
That file is 0 bytes — and it is one of 82 zero-length .drv files. A .drv
can never legitimately be empty. Its same-kernel sibling
(xkygb35z5yw7l4sp2kp0mx260jppd4zc-initrd-linux-6.18.39.drv) is also empty, as is
the whole enclosing closure: nixos-system-machine-test.drv, activate.drv,
etc.drv, initrd-units.drv, boot.json.drv.
Widened sweep: ~4,207 zero-length files across 168 registered store paths.
nix-store --verify --check-contents has now run to completion and confirms that
number by a second, independent method:
| Result | Count |
|---|---|
| Paths reported modified | 168 |
| …truncated to empty (NAR hash of a zero-byte file) | 162 |
| …non-empty hash mismatch | 6 |
The 6 non-empty mismatches are not separate corruption. Every one is a
directory store path (release-notes, three source trees, pkgs-lib, lib)
whose hash changed only because truncated files sit inside it — and every
truncated file inside them carries the same ctime as the rest. One of them,
kjgdxcjz…-source, accounts for 2,845 of the ~4,207 truncated files by itself.
100% of the damage Nix can detect traces to a single minute. No second fault, no ongoing rot.
The damaged set is the NixOS VM-test closure, so this breaks at least six
lanes, not just home-manager:
vm-test-run-kolu-{service,offline-provision,adoption-upgrade,adoption-upgrade-reboot},
kolu-upgrade-verify-reboot, kolu-agent-closure-containment, plus
unit-home-manager-alice.service.drv.
Every truncated file shares one ctime minute, and last -x records that boot
session ending in crash (Sat Aug 22 14:02 - crash):
| Minute | Files created | Zero-length | Rate |
|---|---|---|---|
| 2026-08-22 16:38 | 10,168 | 4,207 | 41.4% |
| 2026-09-05 14:07 | 100,202 | 899 | 0.9% |
| typical minute | ~11,000 | ~19 | 0.2% |
Baseline 0.1–0.9% is legitimately-empty files inside packages. That one minute is 50–400× baseline. Metadata committed, data extents never did — textbook unclean shutdown.
- no I/O errors in
dmesg; btrfs mountedrw, no read-only remount /nix15% used, 1.6 T free- structural
nix-store --verifyclean (it is contents that are lost, not links) - every
.drvwritten since — Aug 24, Sep 3–6, thousands — is healthy - only 1 of 93 flagged paths is a non-empty hash mismatch, and it is a directory containing the same truncated files — same event, not separate bit-rot
This is an Incus container on Nix's experimental local-overlay store. All 82
truncated .drvs are in the container's writable upper layer
(/var/lib/nix/overlay/upper); the shared lower store has zero truncated
.drvs. All six sibling pool boxes (kolu-ci-5,6,7,8,9,10) were checked and are
clean. The base image is intact; the pool is not contaminated.
- Pull
kolu-ci-3from the venue pool immediately. It is still live and still writing (store activity observed the same day), so it keeps handing red lanes to whoever draws it — and the failures read like repo problems. - Recycle the container. It is a disposable
pubox on a verifiably clean lower store, so re-creating it is cheaper and more certain than repairing. - If you would rather repair in place: delete all ~168 damaged paths and let Nix
re-realise them.
nix-store --repairgenuinely cannot help — these are locally-built.drvs and a.drvhas no substituter to re-fetch from. - No disk action needed. If you want belt-and-braces, a btrfs scrub will confirm there is nothing structurally wrong.
Nothing on the host was changed — no repair, no delete, no config or service
touched. sudo requires an interactive password there, and the scope finding
above made the originally-proposed single-file deletion pointless anyway. The
completed verify log is left at /tmp/verify.log for review; temporary scripts
were removed.
ci-home-manager-tail.log— the failing lane's log, including the errorci-home-manager-head.log— the start, showing it got no further than flake input resolutionerror-lines.txt— the two error lines on their own
Seen on juspay/kolu PR #2234, commits 7fe603e25 and c02228625, lane ci::home-manager.