You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
skill-postmortem — a Claude Code skill that turns one observed skill failure into a persistent memory pattern, one gated skill edit, and a ledger entry. Adapted from WikiSkill (arXiv:2608.27454).
A Claude Code skill. Turns one observed skill failure into three artifacts with
different lifetimes: a memory pattern that persists, a skill edit that may be
reverted tomorrow, and a ledger entry recording which of those happened. The
pattern is written before the edit is proposed, so reverting an edit never
destroys the evidence that motivated it.
Adapted from WikiSkill: Compiling Agent Experience into Persistent Knowledge
for Skill Evolution, Tang, Rashtchian, Ferng, Tomkins, Juan, Vu, 27 Aug 2026.
https://arxiv.org/abs/2608.27454
The three mechanisms borrowed from it, the ones deliberately left behind
(validation gating, the evolution loop), and a caveat on the headline numbers
are in references_why.md.
Layout
Gists are flat, so one path is flattened. On disk:
~/.claude/skills/skill-postmortem/
SKILL.md
PURPOSE.md
references/why.md <- references_why.md in this gist
SKILL.md refers to references/why.md by its on-disk path.
The ledger it writes to (~/.claude/skill-impact.md) is not included here; it
holds workspace-specific entries rather than anything reusable.
Not born from a failure. Built 2026-09-01 from WikiSkill (arXiv:2608.27454v1),
which found that a skill-evolution loop improves substantially when the evidence
behind skill edits is kept as a separate, persistent artifact rather than living
inside the optimization history. The workspace already had the evidence layer
(65 files in the project memory directory, indexed from MEMORY.md). It had no
record of skill edits tried and reverted, and no link from a memory to the skill
it should have changed. This skill supplies both.
Reasoning and the numbers: references/why.md.
Patterns addressed
None yet. The first real postmortem run fills this section in.
Evolution history
2026-09-01 created Initial procedure: ledger read first, pattern written
before any edit is proposed, edit gated by Andy rather than by a score,
outcome recorded on rejection as well as acceptance.
Adapted from WikiSkill (arXiv:2608.27454v1, 27 Aug 2026), which co-evolves agent
skills against a persistent knowledge base and gates each edit on validation
accuracy. The full method is a training loop over labeled train/val/test splits
with an automatic scorer, roughly 960 agent rollouts per model per benchmark.
None of that transplants here. Three of its mechanisms do, and the rules in
SKILL.md are those three.
1. The knowledge is never rolled back, only the procedure is
The paper's own contribution. Skill edits are reverted when validation drops;
the pattern pages that motivated them persist across every iteration regardless
of the verdict. Knowledge compounds while procedure oscillates.
This is why the pattern is written in Phase 3 and the edit only proposed in
Phase 4. Reversing that order means a rejected edit takes its evidence with it,
and the same failure gets rediscovered from scratch later.
2. Rejected proposals are results
Their ablation (Table 3, Gemini-3.5-Flash) gives the Skill Proposer access to
the persistent wiki and average benchmark accuracy moves 48.7% to 63.7%, with
LiveMath 51.3% to 72.6% and SpreadsheetBench 49.9% to 76.6%. A large part of
that is the proposer reading which edits were already tried and rejected, and
not proposing them again. Their agent's workflow puts that read at step 2, ahead
of any diagnosis.
Their ALFWorld case study is the mechanism in miniature. At iteration 0 the
proposer suggests goal-directed-action; validation drops and it is rejected.
The rejection record, not the skill, is what steers iteration 1 toward the
concrete rule that is accepted (Never Return an Item to Its Origin Location),
which is refined again at iteration 4 as new evidence accumulates.
Hence Phase 1 before Phase 2, and Phase 6 recording rejections.
3. Keep the notes out of the executing agent's context
The counterintuitive one. Giving the Inference Agent access to the wiki during
rollouts made things worse: 63.7% to 60.9% average, and LiveMath 72.6% to 64.8%.
Their reading is that the agent starts solving from the notes rather than from
the skill, so the resulting trace no longer says anything about whether the
skill works.
Practical form of that here: the notes belong to whoever writes the skill, not
to whoever runs it. SKILL.md stays self-contained and states its rules
outright; PURPOSE.md holds the provenance and is not loaded at execution time.
Smaller borrowings
Both signs of evidence. They sample up to 5 failing and 3 passing traces per
iteration, each capped at 15,000 characters. The passing traces are explicitly
for catching regressions in behavior that already worked. Phase 2 keeps the
shape, drops the ratio.
Index entries carry the fix. They call index.md the most important part of
the wiki, because it alone decides whether a pattern page is ever opened, and
require every entry to state problem, root cause and fix in one or two
sentences. That maps onto the MEMORY.md pointer format in Phase 3.
Atomic proposals, patch preferred. One skill per proposal, patch over create
when the existing skill is partly right, and if most of the file must change,
call it a rewrite instead of a patch.
What was deliberately not taken
Validation gating. The load-bearing quality control in the paper, and it
needs a scorer this workspace does not have for "did the agent work well". Andy
substitutes for it in Phase 5. That is approval, not measurement, and the skill
says so rather than borrowing the paper's language of acceptance.
The evolution loop. Eight iterations of full-batch rollouts with automatic
rescoring. If the loop is ever wanted here it belongs where a pass/fail number
already exists (nightly-maintenance over the fleet: suite results and brakeman
diffs), not bolted onto interactive sessions.
Wiki pruning. They flag its absence as a limitation, since patterns, logs
and diffs accumulate with no automated cleanup. The same will be true of the
ledger. Revisit when it gets long enough to be annoying.
Caveat on the headline numbers
Every model evaluated is small to mid-scale: Qwen-3.5-4B, Qwen-3.5-9B,
Qwen-3.6-27B, Gemma-4-31B, Gemini-3.5-Flash. There is no frontier row. The
no-skill baselines are low (Gemini Flash at 33.0% on LiveMath, 50.5% on
SpreadSheet), so a good share of what evolved skills bought those models is
behavior a frontier model produces unprompted. The paper's trend has gains
rising with capability, but the top of its range is Flash, so applying the
magnitudes above that is extrapolation. The mechanisms above are worth taking on
their logic; the percentages are not a forecast for this workspace.
Use when Andy types /skill-postmortem or says a skill misfired, gave bad guidance, or needs a rule added after it went wrong. Turns one observed skill failure into a durable memory pattern first, then a single gated skill edit, then a ledger entry recording the verdict. Writes the pattern whether or not the edit is accepted, so rejected proposals are never silently re-proposed. Not for authoring a new skill from scratch (that is claude-skill or skill-creator), and not for a plain memory write when no skill was involved.
disable-model-invocation
true
/skill-postmortem — one failure, one pattern, one gated edit
A skill gave bad guidance and you saw it happen. This turns that into three
artifacts with different lifetimes: a pattern that outlives everything, a
skill edit that may be reverted tomorrow, and a ledger entry that
records which of those two happened and why.
The ordering matters more than any individual step. The pattern is written
before the edit is proposed, so that reverting the edit does not destroy the
evidence that motivated it. Everything below exists to protect that ordering.
Method and the numbers behind it: references/why.md.
Scope
One skill per run. One atomic edit per run. If two skills misbehaved, run this
twice.
Do not run this when:
No skill was involved. That is an ordinary memory write, not a postmortem.
The failure was a one-off environment problem (network, a host down, a
half-applied migration). Patterns need evidence that generalises.
You are creating a skill from nothing. Use claude-skill or skill-creator.
Where things live
The ledger sits next to the skill root that owns the skill:
Skill location
Ledger
~/.claude/skills/<name>/
~/.claude/skill-impact.md
<repo>/.claude/skills/<name>/
<repo>/.claude/skill-impact.md, committed
Patterns go to the project memory directory for the cwd, in the existing
one-fact-per-file format, indexed from MEMORY.md.
PURPOSE.md sits inside the skill folder beside SKILL.md. It is provenance,
never instructions, and is not read at execution time.
Phase 1 — Read the ledger before analysing anything
Open the ledger and read every entry for this skill, REJECTED ones first.
This comes first, not last. A rejected entry is a result: someone already tried
that edit and it did not hold. Knowing that before you diagnose stops you from
walking the same path and arriving pleased with yourself at the same dead end.
If the ledger does not exist yet, create it from the template at the bottom.
Phase 2 — Sample the evidence on both signs
Failing side. Where did the skill's guidance produce the wrong action? Quote
the instruction as written and the action that followed. Both, verbatim.
Passing side. Where did this same skill work, in this session or in a
memory you can cite? At least one.
The passing side is not padding. The edit has to fix the failure without
breaking the path that already worked, and you cannot check that against
evidence you never collected.
Phase 3 — Diagnose the root cause, then write the pattern
State why the guidance failed, not what the failure looked like. "The skill said
X, the agent did Y" is a symptom. "The skill said X, which is ambiguous when Z
holds, and Z holds on every repo with a worktree" is a cause.
Write the pattern to memory now, before proposing anything:
One file, one fact, existing frontmatter format.
type: feedback for how work should be done, type: reference for a fact
about a tool or system.
Body carries the root cause and the concrete fix, with [[links]] to related
memories.
Then add the MEMORY.md pointer, and make it carry problem, root cause, and
fix in one line, not just the topic. The pointer is the only thing read when
deciding whether to open the file, so a line naming only the subject is a file
nobody opens.
✅ - [bundle --version lies](ref.md) — reports the nearest lockfile's BUNDLED WITH,
not what's installed; use `gem list bundler --local`
❌ - [bundler versions](ref.md) — notes on bundler version reporting
The pattern is now permanent. Nothing later in this procedure removes it.
Phase 4 — Propose exactly one atomic edit
Prefer patching an existing skill over writing a new one. If the existing skill
is partly right, patch it; a second skill covering adjacent ground splits the
trigger surface and both get worse.
Keep the changed span small and specific. If most of the file has to move, that
is a rewrite, say so and treat it as one rather than dressing it as a patch.
Show the diff. Do not apply it.
Phase 5 — Andy gates it
There is no validation split here and no score. The upstream method accepts an
edit only when measured accuracy rises; nothing in this workspace produces that
number for "did the agent work well". Andy is the gate, so the diff goes to him
unapplied and he accepts, rejects, or amends it.
Never describe the result of this step as verified or measured. It is approved.
Phase 6 — Record the outcome either way
Append to the ledger whether the edit was accepted or rejected. Rejections
are the entries with the most value later, and they are the ones there is a
temptation to skip.
If accepted: apply the edit, then update the skill's PURPOSE.md with the
pattern link and a dated history line. Create PURPOSE.md if absent.
If rejected: the memory pattern still stands from Phase 3. Only the edit is
discarded.
Hard rules
Never write a memory path or [[link]] into a SKILL.md body. A skill is
read at execution time by an agent that should be acting from the procedure,
not reasoning from the research notes behind it. Feeding it both measurably
degrades the result (references/why.md). Provenance belongs in PURPOSE.md.
Never edit a skill before its pattern file exists. That is the one
ordering this whole procedure is built to enforce.
Never re-propose a ledger-rejected edit unless new evidence explains why
the rejection was wrong, and say so in the proposal.
Never state a skill rule the evidence does not support. One trace is an
anecdote. Say "seen once" in PURPOSE.md when that is what happened.
Templates
PURPOSE.md:
# Purpose — <skill-name>## Origin
<why this skill exists, one or two lines>
## Patterns addressed-[[pattern-slug]] — problem + root cause + fix, one line
## Evolution history- YYYY-MM-DD accepted <rule added, and what it fixes>
- YYYY-MM-DD rejected <what was tried, why it did not hold>
Ledger entry, newest at top under the header:
## YYYY-MM-DD <skill-name> ACCEPTED | REJECTED
Pattern: [[pattern-slug]]
Evidence: <failing case, one line> / <passing case, one line>
Change:
<unifieddiffoftheproposededit>
Verdict: <what Andy said, or what broke>
New ledger file header:
# Skill impact ledger
Append-only record of skill edits proposed via /skill-postmortem, accepted and
rejected alike. Read the REJECTED entries before proposing anything: they are
results, not noise.