Skip to content

Instantly share code, notes, and snippets.

View srid's full-sized avatar

Sridhar Ratnakumar srid

View GitHub Profile
@srid
srid / README.md
Last active September 6, 2026 20:43
kolu-ci-3: zero-length .drv in the Nix store (filesystem corruption) — juspay/kolu#2234

kolu-ci-3: ~4,200 truncated files from one crash — recycle the container

Host: kolu-ci-3 (kolu CI venue pool, x86_64-linux, Incus container) Status: still live and taking CI work — pull it from the pool now Verdict: one unclean shutdown two weeks ago, not a failing disk. Recycle the container.

Correction: the first version of this gist said "one specific store path is damaged." That was wrong about scope, and acting on it would not have fixed anything — see below. The symptom and the host are right; the blast radius was understated by three orders of magnitude.

@srid
srid / README.md
Last active September 3, 2026 07:01
olai PR #485 — evidence for the five missing Workflowy chords: the shot driver, its argv, and hashes of the five frames (gists are text-only; the PNGs reproduce in a minute)

PR #485 — the five missing Workflowy chords, on a real serve

The e2e pins are in the PR (features/the_missing_chords.feature, eight scenarios, red-first). This gist is the look of it: five screenshots taken by driver.ts against a dev server reading a copy of the suite's good corpus, the chords pressed by a headless Chromium rather than narrated. The PNGs themselves render on the PR — gists are text-only, so they live at juspay/olai#485 (comment) and are mirrored below via imgur.

@srid
srid / README.md
Created September 2, 2026 20:33
blank-drafts evidence

blank-drafts evidence shots

@srid
srid / debug-macos-xyne-boxes.sh
Created August 17, 2026 23:40
Debug xyne-boxes Killed: 9 on macOS Apple Silicon
#!/bin/sh
# https://github.com/juspay/xyne-boxes/pull/24
# Paste output back after: curl -fsSL https://raw.githubusercontent.com/juspay/xyne-boxes/nightly/installer/install.sh | sh
set -eu
bin="${XYNE_BOXES_BIN:-$HOME/.local/bin}/xyne-boxes"
echo "=== uname ==="
uname -a
echo "=== file ==="
@srid
srid / Orchestrator.md
Last active August 15, 2026 14:20
Orchestrator.md

You are the agent orchestrator for this repository in $PWD. You are responsible for managing multiple tasks, each working in their own toplevel Kolu terminal in their own worktree under $PWD/.worktrees/.

You are expected to be running on a superior model that is also expensive (e.g.: Fable). Therefore, when you spawn subagents, reserve that model (Fable) only where that level of intelligence is necessary.

Planning & Roadmap updates

Have a conversation with the user to flesh out any idea. Use AskUserQuestion where appropriate. Once ready: update the Olai roadmap (after using AskUserQuestion to resolve all ambiguities). All work items have a correspoding Olai roadmap entry. Olai roadmap is kept up to date in $PWD.

The Olai roadmap is written ONLY through olai's own ops (the MCP tools) — never by editing the roadmap file directly, never by jq, never as a git commit the orchestrator authors. Every op validates the whole set; the ops layer is the ledger's only committer. Serve with --commit=manu

@srid
srid / pu-saturation-gist.md
Created July 6, 2026 12:49
pu/Incus cluster saturation — overnight 2026-07-05→06 (kolu #1204)

pu / Incus cluster saturation — overnight 2026-07-05→06 (W4 CI run)

For the pu admin. During an overnight autonomous run (kolu W4 "the switch"), pu create for a second machine (a remote host to prove kolu's cross-machine live-switch) failed repeatedly for ~3 hours with an Incus cluster capacity error, then recovered by morning. Filing what I saw so you can judge whether the per-member cap or the stale-box population needs attention.

Symptom

pu create w4-remote failed on every attempt from 05:35 to 08:06 (11 attempts, ~15 min apart), each dying at the Launching … step. The error (per the cluster) was a capacity rejection — no Incus cluster member had a free slot under the limit: 4 per member cap. A retry loop logged:

05:35 attempt 1: Launching w4-remote
@srid
srid / 403-gist.md
Created July 2, 2026 10:07
kolu CI: GitHub 403 API rate-limit on shared pool NAT IP 219.65.110.2 (unauthenticated 60/hr) — admin

kolu CI: GitHub HTTP 403 "API rate limit exceeded" on the shared pool NAT IP

The kolu-ci-* Incus pool boxes all egress through one shared NAT IP: 219.65.110.2. Unauthenticated requests from that IP to api.github.com now return HTTP 403 — rate limit exceeded (GitHub's unauthenticated limit is 60 requests/hour per IP, and the shared IP has burned all 60).

Raw evidence (from inside kolu-ci-1)

$ curl -s -D - -o /dev/null https://api.github.com
HTTP/2 403
@srid
srid / pu-gist.md
Created July 2, 2026 10:06
kolu CI: lease.sh false-negative 'no egress' (unauth api.github.com rate-limit on shared NAT) — not an outage

kolu CI linux pool — lease.sh false-negative "no egress" (NOT an outage)

TL;DR — the pool boxes are healthy; the egress probe is wrong. ci/pu/lease.sh probes egress with an unauthenticated curl -sf https://api.github.com. GitHub's unauthenticated limit is 60 requests/hour per IP, and all 8 kolu-ci-* boxes share one NAT egress IP (219.65.110.2), which has exhausted that budget → every probe gets HTTP 403 "API rate limit exceeded" → curl -sf (fails on ≥400) exits non-zero → lease.sh marks every box NOEGRESS and skips it. The cold-ephemeral fallback shares the same NAT IP, so it "fails" identically.

Real connectivity from inside kolu-ci-1 is fine:

github.com          -> HTTP 200      (git transport host — what odu's lane fetches)
cache.nixos.org     -> HTTP 200      (nix substituter)
api.github.com      -> HTTP 403      x-ratelimit-limit: 60, x-ratelimit-remaining: 0
@srid
srid / wellbeing.md
Created June 27, 2026 02:18
Debate (Claude × Codex): The most pressing problem to solve first for individual happiness

Debate: What's the most pressing problem we must solve first to ensure individual happiness for every human being?

A back-and-forth between Claude (terminal 7bfb4a4f) and Codex (terminal 53b7f178). Goal: converge on the single MOST pressing problem humanity must solve first so that each and every individual can be happy.


Claude — Opening Position (Round 1)

My claim: the most pressing problem is the systematic erosion of secure belonging — the loss of stable, trusted relationships and community in which a person is known, needed, and not alone.

@srid
srid / 1problem.md
Created June 27, 2026 01:59
Debate: What's the most pressing problem in the world that we humans must solve first? (Claude vs Codex)

Debate: What's the most pressing problem in the world that we humans must solve first?

A back-and-forth between Claude (terminal 02a96e99) and Codex (terminal ce14ad9b). Goal: converge on the single MOST pressing problem humanity must solve first.


Claude — Opening Position (Round 1)

My claim: the most pressing problem is our collapsing capacity for collective coordination — the "meta-problem" of being unable to make and keep civilization-scale decisions.