Skip to content

Instantly share code, notes, and snippets.

@andynu
Created September 9, 2026 16:31
Show Gist options
  • Select an option

  • Save andynu/0acd2896dabb9e539a30f3ba85b74db9 to your computer and use it in GitHub Desktop.

Select an option

Save andynu/0acd2896dabb9e539a30f3ba85b74db9 to your computer and use it in GitHub Desktop.
skill-postmortem — a Claude Code skill that turns one observed skill failure into a persistent memory pattern, one gated skill edit, and a ledger entry. Adapted from WikiSkill (arXiv:2608.27454).

skill-postmortem

A Claude Code skill. Turns one observed skill failure into three artifacts with different lifetimes: a memory pattern that persists, a skill edit that may be reverted tomorrow, and a ledger entry recording which of those happened. The pattern is written before the edit is proposed, so reverting an edit never destroys the evidence that motivated it.

Adapted from WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution, Tang, Rashtchian, Ferng, Tomkins, Juan, Vu, 27 Aug 2026. https://arxiv.org/abs/2608.27454

The three mechanisms borrowed from it, the ones deliberately left behind (validation gating, the evolution loop), and a caveat on the headline numbers are in references_why.md.

Layout

Gists are flat, so one path is flattened. On disk:

~/.claude/skills/skill-postmortem/
  SKILL.md
  PURPOSE.md
  references/why.md       <- references_why.md in this gist

SKILL.md refers to references/why.md by its on-disk path.

The ledger it writes to (~/.claude/skill-impact.md) is not included here; it holds workspace-specific entries rather than anything reusable.

Purpose — skill-postmortem

Origin

Not born from a failure. Built 2026-09-01 from WikiSkill (arXiv:2608.27454v1), which found that a skill-evolution loop improves substantially when the evidence behind skill edits is kept as a separate, persistent artifact rather than living inside the optimization history. The workspace already had the evidence layer (65 files in the project memory directory, indexed from MEMORY.md). It had no record of skill edits tried and reverted, and no link from a memory to the skill it should have changed. This skill supplies both.

Reasoning and the numbers: references/why.md.

Patterns addressed

None yet. The first real postmortem run fills this section in.

Evolution history

  • 2026-09-01 created Initial procedure: ledger read first, pattern written before any edit is proposed, edit gated by Andy rather than by a score, outcome recorded on rejection as well as acceptance.

Why the procedure is ordered this way

Adapted from WikiSkill (arXiv:2608.27454v1, 27 Aug 2026), which co-evolves agent skills against a persistent knowledge base and gates each edit on validation accuracy. The full method is a training loop over labeled train/val/test splits with an automatic scorer, roughly 960 agent rollouts per model per benchmark. None of that transplants here. Three of its mechanisms do, and the rules in SKILL.md are those three.

1. The knowledge is never rolled back, only the procedure is

The paper's own contribution. Skill edits are reverted when validation drops; the pattern pages that motivated them persist across every iteration regardless of the verdict. Knowledge compounds while procedure oscillates.

This is why the pattern is written in Phase 3 and the edit only proposed in Phase 4. Reversing that order means a rejected edit takes its evidence with it, and the same failure gets rediscovered from scratch later.

2. Rejected proposals are results

Their ablation (Table 3, Gemini-3.5-Flash) gives the Skill Proposer access to the persistent wiki and average benchmark accuracy moves 48.7% to 63.7%, with LiveMath 51.3% to 72.6% and SpreadsheetBench 49.9% to 76.6%. A large part of that is the proposer reading which edits were already tried and rejected, and not proposing them again. Their agent's workflow puts that read at step 2, ahead of any diagnosis.

Their ALFWorld case study is the mechanism in miniature. At iteration 0 the proposer suggests goal-directed-action; validation drops and it is rejected. The rejection record, not the skill, is what steers iteration 1 toward the concrete rule that is accepted (Never Return an Item to Its Origin Location), which is refined again at iteration 4 as new evidence accumulates.

Hence Phase 1 before Phase 2, and Phase 6 recording rejections.

3. Keep the notes out of the executing agent's context

The counterintuitive one. Giving the Inference Agent access to the wiki during rollouts made things worse: 63.7% to 60.9% average, and LiveMath 72.6% to 64.8%. Their reading is that the agent starts solving from the notes rather than from the skill, so the resulting trace no longer says anything about whether the skill works.

Practical form of that here: the notes belong to whoever writes the skill, not to whoever runs it. SKILL.md stays self-contained and states its rules outright; PURPOSE.md holds the provenance and is not loaded at execution time.

Smaller borrowings

Both signs of evidence. They sample up to 5 failing and 3 passing traces per iteration, each capped at 15,000 characters. The passing traces are explicitly for catching regressions in behavior that already worked. Phase 2 keeps the shape, drops the ratio.

Index entries carry the fix. They call index.md the most important part of the wiki, because it alone decides whether a pattern page is ever opened, and require every entry to state problem, root cause and fix in one or two sentences. That maps onto the MEMORY.md pointer format in Phase 3.

Atomic proposals, patch preferred. One skill per proposal, patch over create when the existing skill is partly right, and if most of the file must change, call it a rewrite instead of a patch.

What was deliberately not taken

Validation gating. The load-bearing quality control in the paper, and it needs a scorer this workspace does not have for "did the agent work well". Andy substitutes for it in Phase 5. That is approval, not measurement, and the skill says so rather than borrowing the paper's language of acceptance.

The evolution loop. Eight iterations of full-batch rollouts with automatic rescoring. If the loop is ever wanted here it belongs where a pass/fail number already exists (nightly-maintenance over the fleet: suite results and brakeman diffs), not bolted onto interactive sessions.

Wiki pruning. They flag its absence as a limitation, since patterns, logs and diffs accumulate with no automated cleanup. The same will be true of the ledger. Revisit when it gets long enough to be annoying.

Caveat on the headline numbers

Every model evaluated is small to mid-scale: Qwen-3.5-4B, Qwen-3.5-9B, Qwen-3.6-27B, Gemma-4-31B, Gemini-3.5-Flash. There is no frontier row. The no-skill baselines are low (Gemini Flash at 33.0% on LiveMath, 50.5% on SpreadSheet), so a good share of what evolved skills bought those models is behavior a frontier model produces unprompted. The paper's trend has gains rising with capability, but the top of its range is Flash, so applying the magnitudes above that is extrapolation. The mechanisms above are worth taking on their logic; the percentages are not a forecast for this workspace.

name skill-postmortem
description Use when Andy types /skill-postmortem or says a skill misfired, gave bad guidance, or needs a rule added after it went wrong. Turns one observed skill failure into a durable memory pattern first, then a single gated skill edit, then a ledger entry recording the verdict. Writes the pattern whether or not the edit is accepted, so rejected proposals are never silently re-proposed. Not for authoring a new skill from scratch (that is claude-skill or skill-creator), and not for a plain memory write when no skill was involved.
disable-model-invocation true

/skill-postmortem — one failure, one pattern, one gated edit

A skill gave bad guidance and you saw it happen. This turns that into three artifacts with different lifetimes: a pattern that outlives everything, a skill edit that may be reverted tomorrow, and a ledger entry that records which of those two happened and why.

The ordering matters more than any individual step. The pattern is written before the edit is proposed, so that reverting the edit does not destroy the evidence that motivated it. Everything below exists to protect that ordering.

Method and the numbers behind it: references/why.md.

Scope

One skill per run. One atomic edit per run. If two skills misbehaved, run this twice.

Do not run this when:

  • No skill was involved. That is an ordinary memory write, not a postmortem.
  • The failure was a one-off environment problem (network, a host down, a half-applied migration). Patterns need evidence that generalises.
  • You are creating a skill from nothing. Use claude-skill or skill-creator.

Where things live

The ledger sits next to the skill root that owns the skill:

Skill location Ledger
~/.claude/skills/<name>/ ~/.claude/skill-impact.md
<repo>/.claude/skills/<name>/ <repo>/.claude/skill-impact.md, committed

Patterns go to the project memory directory for the cwd, in the existing one-fact-per-file format, indexed from MEMORY.md.

PURPOSE.md sits inside the skill folder beside SKILL.md. It is provenance, never instructions, and is not read at execution time.

Phase 1 — Read the ledger before analysing anything

Open the ledger and read every entry for this skill, REJECTED ones first.

This comes first, not last. A rejected entry is a result: someone already tried that edit and it did not hold. Knowing that before you diagnose stops you from walking the same path and arriving pleased with yourself at the same dead end.

If the ledger does not exist yet, create it from the template at the bottom.

Phase 2 — Sample the evidence on both signs

Failing side. Where did the skill's guidance produce the wrong action? Quote the instruction as written and the action that followed. Both, verbatim.

Passing side. Where did this same skill work, in this session or in a memory you can cite? At least one.

The passing side is not padding. The edit has to fix the failure without breaking the path that already worked, and you cannot check that against evidence you never collected.

Phase 3 — Diagnose the root cause, then write the pattern

State why the guidance failed, not what the failure looked like. "The skill said X, the agent did Y" is a symptom. "The skill said X, which is ambiguous when Z holds, and Z holds on every repo with a worktree" is a cause.

Write the pattern to memory now, before proposing anything:

  • One file, one fact, existing frontmatter format.
  • type: feedback for how work should be done, type: reference for a fact about a tool or system.
  • Body carries the root cause and the concrete fix, with [[links]] to related memories.

Then add the MEMORY.md pointer, and make it carry problem, root cause, and fix in one line, not just the topic. The pointer is the only thing read when deciding whether to open the file, so a line naming only the subject is a file nobody opens.

✅ - [bundle --version lies](ref.md) — reports the nearest lockfile's BUNDLED WITH,
     not what's installed; use `gem list bundler --local`
❌ - [bundler versions](ref.md) — notes on bundler version reporting

The pattern is now permanent. Nothing later in this procedure removes it.

Phase 4 — Propose exactly one atomic edit

Prefer patching an existing skill over writing a new one. If the existing skill is partly right, patch it; a second skill covering adjacent ground splits the trigger surface and both get worse.

Keep the changed span small and specific. If most of the file has to move, that is a rewrite, say so and treat it as one rather than dressing it as a patch.

Show the diff. Do not apply it.

Phase 5 — Andy gates it

There is no validation split here and no score. The upstream method accepts an edit only when measured accuracy rises; nothing in this workspace produces that number for "did the agent work well". Andy is the gate, so the diff goes to him unapplied and he accepts, rejects, or amends it.

Never describe the result of this step as verified or measured. It is approved.

Phase 6 — Record the outcome either way

Append to the ledger whether the edit was accepted or rejected. Rejections are the entries with the most value later, and they are the ones there is a temptation to skip.

If accepted: apply the edit, then update the skill's PURPOSE.md with the pattern link and a dated history line. Create PURPOSE.md if absent.

If rejected: the memory pattern still stands from Phase 3. Only the edit is discarded.

Hard rules

  • Never write a memory path or [[link]] into a SKILL.md body. A skill is read at execution time by an agent that should be acting from the procedure, not reasoning from the research notes behind it. Feeding it both measurably degrades the result (references/why.md). Provenance belongs in PURPOSE.md.
  • Never edit a skill before its pattern file exists. That is the one ordering this whole procedure is built to enforce.
  • Never re-propose a ledger-rejected edit unless new evidence explains why the rejection was wrong, and say so in the proposal.
  • Never state a skill rule the evidence does not support. One trace is an anecdote. Say "seen once" in PURPOSE.md when that is what happened.

Templates

PURPOSE.md:

# Purpose — <skill-name>

## Origin
<why this skill exists, one or two lines>

## Patterns addressed
- [[pattern-slug]] — problem + root cause + fix, one line

## Evolution history
- YYYY-MM-DD  accepted  <rule added, and what it fixes>
- YYYY-MM-DD  rejected  <what was tried, why it did not hold>

Ledger entry, newest at top under the header:

## YYYY-MM-DD  <skill-name>  ACCEPTED | REJECTED
Pattern:  [[pattern-slug]]
Evidence: <failing case, one line> / <passing case, one line>
Change:
    <unified diff of the proposed edit>
Verdict:  <what Andy said, or what broke>

New ledger file header:

# Skill impact ledger

Append-only record of skill edits proposed via /skill-postmortem, accepted and
rejected alike. Read the REJECTED entries before proposing anything: they are
results, not noise.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment