Skip to content

Instantly share code, notes, and snippets.

@srid
Created July 2, 2026 10:06
Show Gist options
  • Select an option

  • Save srid/f387bc17831574a93ac36c5975db5619 to your computer and use it in GitHub Desktop.

Select an option

Save srid/f387bc17831574a93ac36c5975db5619 to your computer and use it in GitHub Desktop.
kolu CI: lease.sh false-negative 'no egress' (unauth api.github.com rate-limit on shared NAT) — not an outage

kolu CI linux pool — lease.sh false-negative "no egress" (NOT an outage)

TL;DR — the pool boxes are healthy; the egress probe is wrong. ci/pu/lease.sh probes egress with an unauthenticated curl -sf https://api.github.com. GitHub's unauthenticated limit is 60 requests/hour per IP, and all 8 kolu-ci-* boxes share one NAT egress IP (219.65.110.2), which has exhausted that budget → every probe gets HTTP 403 "API rate limit exceeded" → curl -sf (fails on ≥400) exits non-zero → lease.sh marks every box NOEGRESS and skips it. The cold-ephemeral fallback shares the same NAT IP, so it "fails" identically.

Real connectivity from inside kolu-ci-1 is fine:

github.com          -> HTTP 200      (git transport host — what odu's lane fetches)
cache.nixos.org     -> HTTP 200      (nix substituter)
api.github.com      -> HTTP 403      x-ratelimit-limit: 60, x-ratelimit-remaining: 0
                                      body: "API rate limit exceeded for 219.65.110.2"
DNS (api.github.com)-> 20.207.73.85  (resolves)

So CI itself would run (git fetch + nix build reach their hosts); only the lease gate wrongly rejects the pool.

The probe that misfires

ci/pu/lease.sh:

egress_ok() { dial "$1" 'timeout 12 curl -sf -o /dev/null https://api.github.com' >/dev/null 2>&1; }
# and inline in the background holder:
timeout 12 curl -sf -o /dev/null https://api.github.com || { echo NOEGRESS; exit 8; }

api.github.com unauthenticated is uniquely fragile here: 60/hr per IP, trivially exhausted on a shared NAT. github.com (200) and cache.nixos.org (200) — the hosts CI actually needs — are not rate-limited this way.

Suggested fixes (any one)

  1. Change the probe target to a host CI truly needs and that isn't unauth-rate-limited — e.g. https://github.com or https://cache.nixos.org/nix-cache-info (both return 200 today). Cheapest, most faithful to "can this box run CI?".
  2. Authenticate the probe — curl -sf -H "Authorization: Bearer $GITHUB_TOKEN" https://api.github.com (auth'd limit is 5000/hr), if a token is on the box.
  3. Accept 403-as-reachable — a 403 proves the TCP+TLS+HTTP round-trip completed, i.e. egress works; drop -f and treat any HTTP response as "reachable."

Admin action

  • Short term: the shared NAT IP 219.65.110.2 is just unauth-rate-limited by GitHub; it clears on its own at the reset window (x-ratelimit-reset), or immediately if the probe stops hitting api.github.com unauthenticated.
  • Real fix: repoint/authenticate the lease.sh egress probe (above) so a rate-limited api.github.com can't false-negative the entire pool.

Filed during a /be ship run for PR #1652. Because the boxes are actually healthy, that run pins a pool box directly (the sanctioned warming path) to bring up the linux lane while this probe fix is pending.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment