Skip to content

Instantly share code, notes, and snippets.

@ElijahLynn
Last active June 8, 2026 02:27
Show Gist options
  • Select an option

  • Save ElijahLynn/3c99af04b68f9c38c39a4717d2db1156 to your computer and use it in GitHub Desktop.

Select an option

Save ElijahLynn/3c99af04b68f9c38c39a4717d2db1156 to your computer and use it in GitHub Desktop.
Ralph Wiggum Loop: Claude Code xhigh vs low effort comparison

Ralph Wiggum Loop: xhigh vs low effort comparison

Testing the Claude Code agentic loop (claude -p --dangerously-skip-permissions) building a small Python CLI from scratch, with two effort levels compared.

Setup

  • App: Ralph Wiggum Quotes CLI — random, list, add, delete subcommands, stdlib only
  • Loop: while grep -q '\- \[ \]' implementation_plan.md — one task per iteration, auto-terminates
  • Branches: main (xhigh/default effort) vs no-thinking (--effort low)
  • Tasks: 8 checkboxes in implementation_plan.md, each committed separately

Speed

Branch Effort Total Per-task avg
main xhigh (default) 451s (7.5 min) 56s
no-thinking low 358s (6 min) 45s

Low effort was 93 seconds faster — ~21% speedup.

Test Coverage

Branch Final test count
main (xhigh) 45 tests
no-thinking (low) 25 tests

Nearly 2x more tests with high effort. The gap grew task-by-task — xhigh consistently wrote 6–8 tests per task, low effort wrote 2–4.

Per-task timing

Task xhigh low effort
1. Data layer 56s 41s
2. Seed data 53s 62s
3. random command 67s 42s
4. list command 47s 38s
5. add command 53s 50s
6. delete command 60s 34s
7. CLI wiring 61s 48s
8. README 54s 43s
Total 451s 358s

Quality differences

Task 5 (add command)

  • xhigh: 8 tests — whitespace stripping, empty-input rejection to stderr, missing-file handling, no-reuse-of-deleted-IDs, persistence, printed ID
  • low: 4 tests — basic increment, persistence, auto-increment, no-reuse

Task 6 (delete command)

  • xhigh: 8 tests — stderr-vs-stdout routing, nonzero exit code, missing-file-as-empty, selective removal, confirmation output
  • low: 3 tests — happy path, missing ID, empty store

Task 7 (CLI wiring)

  • xhigh: 8 tests — int-ID parsing, required-text, persistence, random, delete, add→list round-trip
  • low: 5 tests — add→list round-trip, empty-store random, delete persistence, missing-ID stderr+exit-1, no-command error

Verdict

21% faster with low effort, but ~45% fewer tests and meaningfully less edge-case coverage. For a quick prototype the low-effort tradeoff is fine; for anything production-bound, the xhigh test depth is worth the extra time.

Implementation Plan: Ralph Wiggum Quotes CLI

Tasks are ordered by dependency. Pick the highest-leverage unchecked task each loop iteration.

Tasks

  • 1. Data layer — implement load_quotes(path) and save_quotes(path, quotes) in quotes.py; handle missing file gracefully (return empty list)
  • 2. Seed data — create quotes.json with at least 5 Ralph Wiggum quotes
  • 3. random command — implement and wire up; handle empty store
  • 4. list command — implement and wire up; handle empty store
  • 5. add command — implement with auto-increment ID; persist to quotes.json
  • 6. delete command — implement with ID lookup; print error to stderr if not found
  • 7. CLI wiring — set up argparse with all four subcommands as the main() entry point
  • 8. README — create README.md covering: what the app does, usage examples for all four commands, how to run the tests
=== ralph wiggum loop — Sun Jun 7 18:44:45 PDT 2026 ===
=== iteration 1 ===
Done.
**[task 1] Data layer** — implemented `load_quotes(path)` and `save_quotes(path, quotes)` in `quotes.py`, handling a missing file by returning an empty list. Added 3 unit tests (missing-file, save/load roundtrip, valid JSON output) — all passing. Marked the task checked and committed as `6f3bb51`.
Stopping here as instructed.
time: 41s
=== iteration 2 ===
Done.
**[task 2] Seed data** — complete.
- Created `quotes.json` with 6 Ralph Wiggum quotes (auto-incrementing integer IDs).
- Added `TestSeedData` to `test_quotes.py` verifying the file exists, has ≥5 quotes, and that all entries have unique integer IDs and non-empty text. All 6 tests pass.
- Marked Task 2 checked in `implementation_plan.md` and committed (`aaf96ba`). Not pushed.
Stopping here as instructed.
time: 62s
=== iteration 3 ===
Task 3 complete and committed.
- Added `random_quote(quotes)` to `quotes.py` — returns a random quote's text, or `None` when the store is empty.
- Added `TestRandomCommand` with two tests (random selection stays within the set; empty store returns `None`). All 8 tests pass.
- Checked off Task 3 in `implementation_plan.md` and committed.
Stopping here as instructed — not continuing to the next task.
time: 42s
=== iteration 4 ===
Done. Task 4 (`list` command) is implemented in `quotes.py` via `format_list()`, tested with two passing unit tests, marked complete in the plan, and committed (`47486d2`).
Stopping here as instructed — not continuing to the next task.
time: 38s
=== iteration 5 ===
Done. Task 5 (`add` command) is implemented with auto-increment IDs that don't reuse deleted ones, verified by 4 passing tests, marked complete in the plan, and committed (`28a1f12`). Stopping here.
time: 50s
=== iteration 6 ===
Done.
- **[task 6] delete command** — added `delete_quote(quotes, quote_id)` to `quotes.py`, which removes the matching quote in place and returns `True`/`False` so the caller can report an error to stderr on a missing ID.
- Added `TestDeleteCommand` (existing ID, missing ID, empty store) — all 3 pass.
- Marked Task 6 checked in `implementation_plan.md` and committed (`2b4bf5c`).
Stopping here per instructions.
time: 34s
=== iteration 7 ===
Done.
**[task 7] CLI wiring** — implemented `main()` in `quotes.py` with `argparse` subcommands for `random`, `list`, `add`, and `delete`, dispatching to the existing helper functions. The `delete` command exits with code 1 and writes to stderr on a missing ID; `add`/`delete` persist via `save_quotes`.
- Added 5 CLI tests in `TestCLIWiring` (patching `DATA_FILE` to a temp store): add→list roundtrip, empty-store random message, delete persistence, missing-ID stderr+exit-1, and no-command error. All 22 tests pass.
- Marked task 7 checked in `implementation_plan.md` and committed (`508bd7f`). Not pushed.
Stopping here. Task 8 (README) remains unchecked.
time: 48s
=== iteration 8 ===
`[task 8] Starting: README` — done.
Created `README.md` (app description, usage examples for all four commands, missing-file/empty/error behavior, and how to run the tests), added a `TestReadme` test class verifying the README exists and documents all four commands plus test instructions. All 25 tests pass. Marked task 8 checked and committed (`177e95d`).
This was the last unchecked task — all tasks in the plan are now complete.
time: 43s
=== all tasks complete — total: 358s ===
#!/usr/bin/env fish
set max_iter 20
set i 0
set total_start (date +%s)
set log_file loop-(date +%Y%m%d-%H%M%S).log
echo "=== ralph wiggum loop — $(date) ===" | tee $log_file
echo "log: $log_file"
while grep -q '\- \[ \]' implementation_plan.md
set i (math $i + 1)
if test $i -gt $max_iter
echo "" | tee -a $log_file
echo "!!! hit max iterations ($max_iter) — stopping" | tee -a $log_file
exit 1
end
set iter_start (date +%s)
echo "" | tee -a $log_file
echo "=== iteration $i ===" | tee -a $log_file
cat prompt.md | claude -p --dangerously-skip-permissions 2>&1 | tee -a $log_file
set claude_exit $pipestatus[2]
set iter_end (date +%s)
echo " time: "(math $iter_end - $iter_start)"s" | tee -a $log_file
if test $claude_exit -ne 0
echo "!!! claude exited $claude_exit — stopping" | tee -a $log_file
exit 1
end
end
set total_end (date +%s)
echo "" | tee -a $log_file
echo "=== all tasks complete — total: "(math $total_end - $total_start)"s ===" | tee -a $log_file
=== ralph wiggum loop — Sun Jun 7 19:06:06 PDT 2026 ===
=== iteration 1 ===
Task 1 complete and committed.
- **Implemented**: `load_quotes(path)` and `save_quotes(path, quotes)` in `quotes.py` — missing file returns `[]`; saves are indented JSON with unicode preserved.
- **Tested**: 4 passing unit tests (missing-file, round-trip, overwrite, unicode) using a temporary store, not the real `quotes.json`.
- **Marked** Task 1 checked in `implementation_plan.md`.
- **Committed** as `68d20a9` (not pushed).
Stopping here as instructed — not proceeding to Task 2.
time: 56s
=== iteration 2 ===
Done.
`[task 2] Starting: Seed data` — completed.
- Created `quotes.json` with 6 valid Ralph Wiggum quotes (≥5 required), auto-incrementing unique integer IDs.
- Added a `TestSeedData` suite to `test_quotes.py` verifying the file exists, has ≥5 quotes, every quote has an integer `id` and non-empty `text`, and all IDs are unique. All 8 tests pass.
- Marked Task 2 checked in `implementation_plan.md`.
- Committed as `5d74bcc` (not pushed).
Stopping here as instructed — not continuing to the next task.
time: 53s
=== iteration 3 ===
Done.
**`[task 3] Starting: random command`** — completed.
- Added `cmd_random(quotes)` to `quotes.py`: prints a random quote as `[id] text`, or a helpful "no quotes yet" message when the store is empty.
- Added `TestRandomCommand` to `test_quotes.py` with 4 unbiased tests: output is always a valid quote line, every quote is reachable over many runs, the empty store prints a non-quote message, and the empty case doesn't raise.
- All 12 tests pass.
- Marked task 3 checked in `implementation_plan.md`.
- Committed as `bf9bc02`.
Stopping here as instructed — not proceeding to the next task.
time: 67s
=== iteration 4 ===
Task 4 complete.
- **Implemented** `cmd_list(quotes)` in `quotes.py` — prints each quote as `[ID] text`, or a helpful message when the store is empty.
- **Tested** with 4 new unit tests in `TestListCommand` (order preserved, one line per quote, empty-store message, no exception on empty). All 16 tests pass.
- **Marked** task 4 checked in `implementation_plan.md`.
- **Committed** as `0600336`.
Stopping here as instructed — not continuing to the next task.
time: 47s
=== iteration 5 ===
Done. Task 5 (`add` command) is complete.
- Added `add_quote(quotes, text)` (auto-increment ID via `max(...)+1`, so deleted IDs aren't reused) and `cmd_add(path, text)` (strips whitespace, rejects empty input to stderr with exit 1, persists, prints assigned ID) in `quotes.py`.
- Added 8 tests in `TestAddCommand` covering ID-from-empty, auto-increment, no-reuse-of-deleted-IDs, persistence, missing-file handling, printed ID, whitespace stripping, and empty-input rejection.
- Full suite: 24 tests pass.
- Marked Task 5 checked in `implementation_plan.md` and committed as `3ae4168`.
Stopping here as instructed — not continuing to the next task.
time: 53s
=== iteration 6 ===
Done.
**[task 6] delete command** — completed.
- Added `delete_quote(quotes, quote_id)` (pure removal, returns `True`/`False`) and `cmd_delete(path, quote_id)` (loads, deletes, persists, confirms; exits 1 with a stderr message when the ID isn't found), matching the existing function-over-class style and `cmd_add` error pattern.
- Added 8 unbiased tests in `TestDeleteCommand` covering: removal + return value, missing-ID handling, selective removal, disk persistence, confirmation output, nonzero exit on missing ID, stderr-vs-stdout routing, and missing-file-treated-as-empty.
- Full suite: **32 tests pass**.
- Marked Task 6 checked in `implementation_plan.md` and committed (`adebd9f`). Not pushed.
Stopping here as instructed. Tasks 7 (CLI wiring) and 8 (README) remain.
time: 60s
=== iteration 7 ===
Done. Task 7 (CLI wiring) is complete:
- Added `argparse` import and a `build_parser()` + `main()` entry point to `quotes.py`, dispatching all four subcommands (`random`, `list`, `add`, `delete`) and loading/persisting via `DEFAULT_PATH`.
- Added 8 unbiased tests in `TestCliWiring` (subcommand parsing, missing-subcommand exit, add→list round-trip, persistence, random, delete, int-ID parsing, required-text). All 40 tests pass.
- Marked Task 7 checked in `implementation_plan.md`.
- Committed as `a2b84e5` (not pushed).
Stopping here — Task 8 (README) remains for the next iteration.
time: 61s
=== iteration 8 ===
Done. Summary of what I completed:
**[task 8] README**
- Created `README.md` covering what the app does, usage examples for all four commands (`random`, `list`, `add`, `delete`), the data store format, and how to run the tests (both pytest and unittest).
- Added a `TestReadme` class to `test_quotes.py` with 5 tests verifying the README exists, describes the app, documents and shows examples for all four commands, and explains how to run the tests. All pass.
- Full suite: **45 tests, OK**.
- Marked task 8 checked in `implementation_plan.md` and committed (`86b4d20`). Not pushed.
All tasks in the plan are now complete. Stopping here as instructed.
time: 54s
=== all tasks complete — total: 451s ===
#!/usr/bin/env fish
set max_iter 20
set i 0
set total_start (date +%s)
set log_file loop-(date +%Y%m%d-%H%M%S).log
echo "=== ralph wiggum loop — $(date) ===" | tee $log_file
echo "log: $log_file"
while grep -q '\- \[ \]' implementation_plan.md
set i (math $i + 1)
if test $i -gt $max_iter
echo "" | tee -a $log_file
echo "!!! hit max iterations ($max_iter) — stopping" | tee -a $log_file
exit 1
end
set iter_start (date +%s)
echo "" | tee -a $log_file
echo "=== iteration $i ===" | tee -a $log_file
cat prompt.md | claude -p --dangerously-skip-permissions 2>&1 | tee -a $log_file
set claude_exit $pipestatus[2]
set iter_end (date +%s)
echo " time: "(math $iter_end - $iter_start)"s" | tee -a $log_file
if test $claude_exit -ne 0
echo "!!! claude exited $claude_exit — stopping" | tee -a $log_file
exit 1
end
end
set total_end (date +%s)
echo "" | tee -a $log_file
echo "=== all tasks complete — total: "(math $total_end - $total_start)"s ===" | tee -a $log_file
  1. Study spec.md thoroughly.

  2. Study implementation_plan.md thoroughly.

  3. Pick the highest-leverage unchecked task from implementation_plan.md. If there are no unchecked tasks, print "All tasks complete." and stop immediately.

  4. Print a single line: [task N] Starting: <task name> where N is the task number and task name is its title.

  5. Complete ONLY this one task. Do not proceed to any other task.

  6. Write an unbiased unit test in test_quotes.py to verify the task is correctly implemented. Run the test and confirm it passes.

  7. Mark the task as checked in implementation_plan.md.

  8. Commit the changes to the repository (if one does not exist, create it first). Do not push anywhere.

Stop. Do not continue to the next task.

Spec: Ralph Wiggum Quotes CLI

Objective

A small Python CLI that manages a collection of Ralph Wiggum quotes. The primary purpose is to serve as a well-scoped test target for the Ralph Wiggum agentic loop — simple enough to complete in ~7 loop iterations, complex enough to exercise real implementation decisions.

Commands

Command Description
python quotes.py random Print a random quote
python quotes.py list Print all quotes with their IDs
python quotes.py add "quote text" Add a new quote, print assigned ID
python quotes.py delete <id> Delete quote by ID, confirm deletion

Project Structure

sandbox/
├── spec.md
├── implementation_plan.md
├── prompt.md
├── quotes.py          # CLI entry point + all logic
├── quotes.json        # persistent data store (auto-created if missing)
└── test_quotes.py     # unit tests

Data Format

quotes.json is a JSON array of objects:

[
  { "id": 1, "text": "I bent my wookie." },
  { "id": 2, "text": "My cat's breath smells like cat food." }
]

IDs are auto-incrementing integers. Deleted IDs are not reused.

Code Style

  • Single file (quotes.py) — no unnecessary modules or abstractions
  • Standard library only: json, argparse, random, pathlib
  • Functions over classes; keep each function under 20 lines
  • Exit with code 1 and a clear message on any error

Testing Strategy

  • Unit tests in test_quotes.py using unittest
  • Tests operate on a temporary in-memory data store (not the real quotes.json)
  • Each task in implementation_plan.md gets at least one test
  • Run with: python -m pytest test_quotes.py -v or python -m unittest test_quotes.py

Acceptance Criteria

  • random prints a quote when quotes exist; prints a helpful message when empty
  • list prints all quotes formatted as [ID] quote text; prints a message when empty
  • add persists the quote and prints the new ID
  • delete removes the quote by ID; prints an error if ID not found
  • All commands handle a missing quotes.json gracefully (treat as empty)
  • All acceptance criteria have a passing unit test

Boundaries

Always do:

  • Validate user input before operating on data
  • Print actionable error messages to stderr

Never do:

  • Depend on third-party packages
  • Modify spec.md or implementation_plan.md during implementation
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment