Skip to content

Instantly share code, notes, and snippets.

@karpathy
Created April 4, 2026 16:25
Show Gist options
  • Select an option

  • Save karpathy/442a6bf555914893e9891c11519de94f to your computer and use it in GitHub Desktop.

Select an option

Save karpathy/442a6bf555914893e9891c11519de94f to your computer and use it in GitHub Desktop.
llm-wiki

LLM Wiki

A pattern for building personal knowledge bases using LLMs.

This is an idea file, it is designed to be copy pasted to your own LLM Agent (e.g. OpenAI Codex, Claude Code, OpenCode / Pi, or etc.). Its goal is to communicate the high level idea, but your agent will build out the specifics in collaboration with you.

The core idea

Most people's experience with LLMs and documents looks like RAG: you upload a collection of files, the LLM retrieves relevant chunks at query time, and generates an answer. This works, but the LLM is rediscovering knowledge from scratch on every question. There's no accumulation. Ask a subtle question that requires synthesizing five documents, and the LLM has to find and piece together the relevant fragments every time. Nothing is built up. NotebookLM, ChatGPT file uploads, and most RAG systems work this way.

The idea here is different. Instead of just retrieving from raw documents at query time, the LLM incrementally builds and maintains a persistent wiki — a structured, interlinked collection of markdown files that sits between you and the raw sources. When you add a new source, the LLM doesn't just index it for later retrieval. It reads it, extracts the key information, and integrates it into the existing wiki — updating entity pages, revising topic summaries, noting where new data contradicts old claims, strengthening or challenging the evolving synthesis. The knowledge is compiled once and then kept current, not re-derived on every query.

This is the key difference: the wiki is a persistent, compounding artifact. The cross-references are already there. The contradictions have already been flagged. The synthesis already reflects everything you've read. The wiki keeps getting richer with every source you add and every question you ask.

You never (or rarely) write the wiki yourself — the LLM writes and maintains all of it. You're in charge of sourcing, exploration, and asking the right questions. The LLM does all the grunt work — the summarizing, cross-referencing, filing, and bookkeeping that makes a knowledge base actually useful over time. In practice, I have the LLM agent open on one side and Obsidian open on the other. The LLM makes edits based on our conversation, and I browse the results in real time — following links, checking the graph view, reading the updated pages. Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase.

This can apply to a lot of different contexts. A few examples:

  • Personal: tracking your own goals, health, psychology, self-improvement — filing journal entries, articles, podcast notes, and building up a structured picture of yourself over time.
  • Research: going deep on a topic over weeks or months — reading papers, articles, reports, and incrementally building a comprehensive wiki with an evolving thesis.
  • Reading a book: filing each chapter as you go, building out pages for characters, themes, plot threads, and how they connect. By the end you have a rich companion wiki. Think of fan wikis like Tolkien Gateway — thousands of interlinked pages covering characters, places, events, languages, built by a community of volunteers over years. You could build something like that personally as you read, with the LLM doing all the cross-referencing and maintenance.
  • Business/team: an internal wiki maintained by LLMs, fed by Slack threads, meeting transcripts, project documents, customer calls. Possibly with humans in the loop reviewing updates. The wiki stays current because the LLM does the maintenance that no one on the team wants to do.
  • Competitive analysis, due diligence, trip planning, course notes, hobby deep-dives — anything where you're accumulating knowledge over time and want it organized rather than scattered.

Architecture

There are three layers:

Raw sources — your curated collection of source documents. Articles, papers, images, data files. These are immutable — the LLM reads from them but never modifies them. This is your source of truth.

The wiki — a directory of LLM-generated markdown files. Summaries, entity pages, concept pages, comparisons, an overview, a synthesis. The LLM owns this layer entirely. It creates pages, updates them when new sources arrive, maintains cross-references, and keeps everything consistent. You read it; the LLM writes it.

The schema — a document (e.g. CLAUDE.md for Claude Code or AGENTS.md for Codex) that tells the LLM how the wiki is structured, what the conventions are, and what workflows to follow when ingesting sources, answering questions, or maintaining the wiki. This is the key configuration file — it's what makes the LLM a disciplined wiki maintainer rather than a generic chatbot. You and the LLM co-evolve this over time as you figure out what works for your domain.

Operations

Ingest. You drop a new source into the raw collection and tell the LLM to process it. An example flow: the LLM reads the source, discusses key takeaways with you, writes a summary page in the wiki, updates the index, updates relevant entity and concept pages across the wiki, and appends an entry to the log. A single source might touch 10-15 wiki pages. Personally I prefer to ingest sources one at a time and stay involved — I read the summaries, check the updates, and guide the LLM on what to emphasize. But you could also batch-ingest many sources at once with less supervision. It's up to you to develop the workflow that fits your style and document it in the schema for future sessions.

Query. You ask questions against the wiki. The LLM searches for relevant pages, reads them, and synthesizes an answer with citations. Answers can take different forms depending on the question — a markdown page, a comparison table, a slide deck (Marp), a chart (matplotlib), a canvas. The important insight: good answers can be filed back into the wiki as new pages. A comparison you asked for, an analysis, a connection you discovered — these are valuable and shouldn't disappear into chat history. This way your explorations compound in the knowledge base just like ingested sources do.

Lint. Periodically, ask the LLM to health-check the wiki. Look for: contradictions between pages, stale claims that newer sources have superseded, orphan pages with no inbound links, important concepts mentioned but lacking their own page, missing cross-references, data gaps that could be filled with a web search. The LLM is good at suggesting new questions to investigate and new sources to look for. This keeps the wiki healthy as it grows.

Indexing and logging

Two special files help the LLM (and you) navigate the wiki as it grows. They serve different purposes:

index.md is content-oriented. It's a catalog of everything in the wiki — each page listed with a link, a one-line summary, and optionally metadata like date or source count. Organized by category (entities, concepts, sources, etc.). The LLM updates it on every ingest. When answering a query, the LLM reads the index first to find relevant pages, then drills into them. This works surprisingly well at moderate scale (~100 sources, ~hundreds of pages) and avoids the need for embedding-based RAG infrastructure.

log.md is chronological. It's an append-only record of what happened and when — ingests, queries, lint passes. A useful tip: if each entry starts with a consistent prefix (e.g. ## [2026-04-02] ingest | Article Title), the log becomes parseable with simple unix tools — grep "^## \[" log.md | tail -5 gives you the last 5 entries. The log gives you a timeline of the wiki's evolution and helps the LLM understand what's been done recently.

Optional: CLI tools

At some point you may want to build small tools that help the LLM operate on the wiki more efficiently. A search engine over the wiki pages is the most obvious one — at small scale the index file is enough, but as the wiki grows you want proper search. qmd is a good option: it's a local search engine for markdown files with hybrid BM25/vector search and LLM re-ranking, all on-device. It has both a CLI (so the LLM can shell out to it) and an MCP server (so the LLM can use it as a native tool). You could also build something simpler yourself — the LLM can help you vibe-code a naive search script as the need arises.

Tips and tricks

  • Obsidian Web Clipper is a browser extension that converts web articles to markdown. Very useful for quickly getting sources into your raw collection.
  • Download images locally. In Obsidian Settings → Files and links, set "Attachment folder path" to a fixed directory (e.g. raw/assets/). Then in Settings → Hotkeys, search for "Download" to find "Download attachments for current file" and bind it to a hotkey (e.g. Ctrl+Shift+D). After clipping an article, hit the hotkey and all images get downloaded to local disk. This is optional but useful — it lets the LLM view and reference images directly instead of relying on URLs that may break. Note that LLMs can't natively read markdown with inline images in one pass — the workaround is to have the LLM read the text first, then view some or all of the referenced images separately to gain additional context. It's a bit clunky but works well enough.
  • Obsidian's graph view is the best way to see the shape of your wiki — what's connected to what, which pages are hubs, which are orphans.
  • Marp is a markdown-based slide deck format. Obsidian has a plugin for it. Useful for generating presentations directly from wiki content.
  • Dataview is an Obsidian plugin that runs queries over page frontmatter. If your LLM adds YAML frontmatter to wiki pages (tags, dates, source counts), Dataview can generate dynamic tables and lists.
  • The wiki is just a git repo of markdown files. You get version history, branching, and collaboration for free.

Why this works

The tedious part of maintaining a knowledge base is not the reading or the thinking — it's the bookkeeping. Updating cross-references, keeping summaries current, noting when new data contradicts old claims, maintaining consistency across dozens of pages. Humans abandon wikis because the maintenance burden grows faster than the value. LLMs don't get bored, don't forget to update a cross-reference, and can touch 15 files in one pass. The wiki stays maintained because the cost of maintenance is near zero.

The human's job is to curate sources, direct the analysis, ask good questions, and think about what it all means. The LLM's job is everything else.

The idea is related in spirit to Vannevar Bush's Memex (1945) — a personal, curated knowledge store with associative trails between documents. Bush's vision was closer to this than to what the web became: private, actively curated, with the connections between documents as valuable as the documents themselves. The part he couldn't solve was who does the maintenance. The LLM handles that.

Note

This document is intentionally abstract. It describes the idea, not a specific implementation. The exact directory structure, the schema conventions, the page formats, the tooling — all of that will depend on your domain, your preferences, and your LLM of choice. Everything mentioned above is optional and modular — pick what's useful, ignore what isn't. For example: your sources might be text-only, so you don't need image handling at all. Your wiki might be small enough that the index file is all you need, no search engine required. You might not care about slide decks and just want markdown pages. You might want a completely different set of output formats. The right way to use this is to share it with your LLM agent and work together to instantiate a version that fits your needs. The document's only job is to communicate the pattern. Your LLM can figure out the rest.

@FERRSTUDIO

Copy link
Copy Markdown

Shameless plug for my implementation, once again(再来安利一次我的实现):https://github.com/ChavesLiu/second-brain-skill
No it's not, and you delu.
The one from Karpathy is a feather.

@gowtham0992

Copy link
Copy Markdown

Link 3.0 is out

Link is local memory for AI agents: plain Markdown files, review-gated writes, no LLM in the memory layer, one store shared by Claude Code, Codex, Cursor, Kiro, Windsurf, Zed, VS Code, Copilot, and Gemini.

What's new:

  • Recall works outside English. The tokenizer split on [^a-z0-9]+, so every non-Latin script produced zero tokens: Japanese, Chinese, Korean, Russian, Arabic and Indic memories were unfindable, and nothing said so. Accented Latin was little better: "déploiement" could not be found by typing "deploiement". Scripts written without spaces are now cut into character bigrams (the Lucene CJKAnalyzer approach, no dictionary), Latin accents fold, and combining marks in Indic scripts are kept because they are vowels, not accents; stripping them turned "मंगलवार" into "गलव". Wiki full-text search had the same bug one layer down and is fixed the same way. Found while verifying it end to end: page filenames also slugged non-Latin titles to nothing, so every such memory was filed as memory.md and the second one was refused as a duplicate. A non-English user could save exactly one memory. Titles now keep their own script. ASCII text takes the exact old path, so nothing existing changes: all nine LoCoMo figures are unchanged across 1,536 third-party queries.

  • "lnk stale": notice when a memory outlived the code it describes. The most repeated complaint about agent memory is that nothing tells you when a memory stopped being true. A note says the parser lives in a/b.py, the file is renamed, the memory keeps being retrieved and believed. Hosted memory services cannot fix this because they never see the repository. Run "lnk stale" inside a repo and it lists memories naming files git no longer has, with the successor path where git recorded a rename. A path is questioned only when it is missing now and git tracked it before; without the second half, an unresolvable path is just prose and flagging it is the noise that teaches people to ignore the flag. Read-only, findings go to the review gate. Precision is measured, not asserted: 0 false flags across 108 path references in Link's own docs, every probed deletion detected, and the eval fails CI if either moves. Stale memories are also marked in the recall packet, so the agent is told rather than left to trust.

image
  • The retrieval benchmark reports precision, not only recall. Recall is the number this category publishes, and it cannot separate a system that retrieves cleanly from one that returns everything, because returning everything scores 1.0 by construction. On the same 1,536 LoCoMo queries a whole-store dump carries 0.26% signal; Link's top-1 packet reaches 0.3086 precision on the fast tier, 117x, and pays a real recall cost that is published in the same table. The track now reports the ceiling each cutoff allows (evidence sets average 1.53 turns, so nobody can exceed precision@10 of 0.152) and R-precision as the k-independent figure to compare across systems.

  • Measured and declined: usage-aware ranking. Four formulations were built and measured (additive frequency, tiebreak-only, recency decay in the Generative Agents form across the recommended 7 to 30 day half-life, MMR diversity) and none ship. Every one either made memories that had gone unread harder to find or did nothing; recency was worst, the old half losing 0.0510 while the fresh half gained 0.0204. The reason is a category difference: those policies suit episodic observation streams, and Link stores durable constraints, which do not become less true for going unread. That is when they most need surfacing. The eval stays in the repo so the next attempt has to clear the same bar.

  • "lnk ingest" for structured exports, contributed by @jakobtfaber. Plan-first: provenance manifests hashed per output, staging through a temp directory, validation before promotion, explicit --replace-unmanaged and --prune gates. Imported docs land in the wiki, never in memory, and stay out of personal-memory proposals.

  • LinkBar 1.4. Bugs first: health probes ran on the main thread and froze the popover on every Status refresh; the review inbox showed five items and hid the rest; the workspace could not be changed in the shipped app because a Finder-launched app never sees LINK_WORKSPACE. All fixed: probes run concurrently off the main thread, the inbox scrolls, the workspace is chosen in Settings, every approve/archive/accept confirms what it did, and a refused save reports the CLI's real reason. Then the 3.0 tie-in: a Status row runs "lnk stale" against the repo your last agent session was in, with a one-click filter on the Memory tab and an amber dot in the menu bar.

image
# macOS, CLI + menu bar app
brew install --cask gowtham0992/link/linkbar
lnk setup

# CLI only (or Linux)
brew install gowtham0992/link/link
lnk setup

# already running Link
brew upgrade && lnk setup

# stale check, from inside a repo
lnk stale

Still: every memory a plain file you can open, nothing durable without review, no LLM in the memory layer, CI blocks network code in the runtime.

Release notes: https://github.com/gowtham0992/link/releases/tag/v3.0.0

Repo: https://github.com/gowtham0992/link
Site: https://gowtham0992.github.io/link/
PyPI: https://pypi.org/project/link-mcp/
MCP: https://registry.modelcontextprotocol.io/?q=io.github.gowtham0992%2Flink
Benchmarks: https://github.com/gowtham0992/link/blob/main/benchmarks/RESULTS.md

@chimezie

Copy link
Copy Markdown

This is a very powerful paradigm and architectural style. I mainly use OpenCode and want to implement this for that harness, but I don't want to reinvent the wheel if there are existing OpenCode implementations. Are there any implementations that can easily be 'extended' to do so?

@equationalapplications

Copy link
Copy Markdown

@chimezie I built something along these lines (disclosure: I'm the author). Curated Thoughts is an open-source (MIT) desktop app for Linux, macOS and Windows. It keeps a Karpathy-style LLM wiki over a plain-markdown vault, backed by a local SQLite "brain". It also ships an MCP server, curated-thoughts-mcp, so any harness can use it. With 2.12 the server exposes 16 tools, including wiki_context, wiki_search, wiki_traverse_graph, vault_semantic_search, vault_write_note, and a review-before-merge proposal flow for curated "wisdom" entries.

For OpenCode there's now a packaged integration: opencode v0.1.0.

  1. Install Curated Thoughts from its releases page. The installers include the curated-thoughts-mcp sidecar. Run the app once to create your brain.

  2. Download opencode-0.1.0.tar.gz from the release above, unpack it, and run:

    ./scripts/install.sh                    # preview: prints exactly what it will change, writes nothing
    CT_INSTALL_EDIT=1 ./scripts/install.sh  # apply

    It adds the mcp entry to your global opencode.json without disturbing comments or formatting, installs a small plugin that puts a memory health snapshot in the system prompt, and adds three skills that teach the agent how to use the wiki.

  3. Check it with opencode mcp list (should show curated-thoughts connected), or run the bundled doctor for a full diagnosis of sidecar, brain, vault and registration.

If you'd rather wire it up by hand, the core is just an MCP entry:

{ "mcp": { "curated-thoughts": { "type": "local", "command": ["curated-thoughts-mcp", "--mcp"], "enabled": true } } }

On extending it: the integrations repo is MIT and each integration is small. There are sibling integrations for Hermes and DeepSeek Harness if you want to see how the pieces fit. Issues and PRs are welcome.

@chimezie

Copy link
Copy Markdown

Thank you and for the info and this implementation

@Withylele

Copy link
Copy Markdown

Thank you

@ZeroDot1

Copy link
Copy Markdown

@spasm-myelixlabs

spasm-myelixlabs commented Sep 19, 2026 •

Copy link
Copy Markdown

A similar idea we've been building since January—but for codebases specifically. As a developer, we needed agents to be far more accurate than they are. Re-grepping and repeating the same thing is not only painful to watch, but it also gets expensive in token usage.

We started with a Python implementation. The core was the same insight: stop re-deriving structure from scratch on every query. Build a persistent graph the agent reads instead of grepping. It worked, but Python was too slow for a live in-memory graph. We rewrote it in Elixir — in-memory tables, concurrent indexing, the graph updates as files change on disk in milliseconds.

The part we took further: the wiki is read-only in your pattern. We added write safety — the graph simulates edits in memory, checks every caller, and rolls back if the write would break something, also explaining to the LLM what will happen. The agent doesn't just understand the codebase; it can reason about whether a change is safe before it touches disk. Searches are ultra-fast, one or two queries for full understanding of what is needed.

Appreciate you writing this up. The core idea — knowledge should compound, not be re-derived — is the thing that separates useful agent tools from expensive grep wrappers. Also, it's been a mantra for development for as long as I can remember: DRY - Don't Repeat Yourself.

Please do give it a go; it's free - I would absolutely love your feedback. This is now the precursor to something much bigger we are working on to make agents dream, learn and be better when they wake. No compression, only relevant context for the task at hand.

https://synapse-mcp.dev

Feedback, comments, notes, are welcome please, more info here: https://github.com/myelixlabs/synapse-mcp

@crajah

crajah commented Sep 20, 2026

Copy link
Copy Markdown

@madgodinc is right that conflict resolution belongs at write time, and the
reason is structural rather than stylistic.

A lint pass has to find the contradiction before it can resolve it, and
finding it is the expensive half: two pages that disagree are discoverable only
by reading both and noticing. That is quadratic in pages, it runs on every
sweep, and an LLM adjudicates each comparison. At write time the comparison is
already local & you hold the arriving assertion and the one it lands on. The
cost collapses to a lookup.

The price is that you must declare in advance which claims are mutually
exclusive. We do that with predicate groups: sets of relations that cannot
simultaneously hold between the same ordered pair. On company filings,
{generates_cash_flow, consumes_cash} and {grounds, certifies} are two of
eight. A later document asserting one against a pair that already holds the
other closes the earlier edge. No sweep, no adjudication, no second LLM call.

Two design rules make this work on real corpora.

Order by document, never by extracted dates. The intuitive design asks the
model when each fact was true and sorts on that. It does not survive contact
with source material: models state a period only when the text happens to say
so, which is rare. Document order is always available - publication order for a
book series, fiscal year for filings, ingest order for everything else - and it
is robust to backfilling. Index 2019 filings after 2024 ones and resolution is
still correct, because the declared order governs, not the write order.
@gowtham0992's use of git history to spot stale memories is the same instinct;
the repository already knows which assertion came later, and that ordering is
more reliable than anything a model will tell you about when something was
true.

Close, never delete. A superseded edge is marked and timestamped, and stays
retrievable on request. This is the gist's own "compounding artifact" argument
taken to its conclusion: a wiki that silently overwrites yesterday's claim is
worth less than one that can show you the succession.

It fires on ordinary prose, not just structured filings. On the d'Artagnan
trilogy - three novels, ~645k chars, indexed in publication order - thirteen
relationships were closed by a later book, including the ally_of edge between
d'Artagnan and Aramis once a later volume recasts them as opponents. The
extractor supplied no dates for any of those pairs.

The pattern needs a second clock. "What is true" and "what did this system
believe in March" are different questions with different answers, because a
source published today can restate a figure from 2019. Both need to be
recoverable, and the second is the one that matters when someone asks why a
decision was made. Model them as separate axes - valid time for when a fact
held in the world, belief time for when the system came to hold it - and
reproducing a past answer is a query parameter rather than an archaeology
expedition through old snapshots. Markdown makes this genuinely hard; it is the
strongest argument for structure I know.


What this is worth, on benchmarks built around exactly this problem.

LongMemEval is long-horizon chat memory, and two of its six categories are the
gist's problem directly: knowledge-update asks whether a fact can be
distinguished from the fact that replaced it, and temporal-reasoning asks for
ordering and intervals. Full benchmark, nothing sampled, against Zep/Graphiti's
published figures:

ours (gemini-3.6-flash) Zep (gpt-4o)
temporal-reasoning 96.2% 62.4%
knowledge-update 94.9% 83.3%
multi-session 90.2% 57.9%
overall 94.0% 71.2%

A full-context baseline that simply puts the entire history in the prompt
scores 60.2% - worth keeping in mind whenever someone suggests a long context
window makes this architecture unnecessary.

Read that as informative about the tier rather than a controlled head-to-head:
their figures are as published, judged by a single gpt-4o where ours is a
three-model panel, and two years separate the model generations in a direction
I cannot sign. One instance of 500 is excluded (a session both extraction
prompts refused) so the denominator is 499.

The mechanism behind the temporal categories is isolable. Carrying each
relation's validity period through to synthesis (not merely extracting and
storing it) is worth +38.4 points on temporal-reasoning and +16.3 on
knowledge-update
, ablated paired on one fixed graph per instance so that
re-indexing noise cannot account for it. Storing time is not the hard part.
Rendering it into the prompt is.

ECT-QA is the adversarial version: earnings-call transcripts, six companies,
sixteen consecutive quarters each, where every metric is restated every quarter
so only the date distinguishes sixteen competing values. Under that benchmark's
own element-wise protocol it scores 0.807 Correct against 0.599 for TG-RAG and
0.406/0.405 for LightRAG and GraphRAG as published - on their corpus slice with
a judge we cannot run, so treat a few points of the margin as approximate.

The reason these are worth citing in this thread is not the ranking. It is that
the two categories which separate systems most are the two that turn on time,
and the gap between handling temporal evolution and ignoring it is measurable
at benchmark scale rather than being an architectural preference.


@ShootJackal's "model output as untrusted input" is the right frame, and it
has a stricter form on the write path than on the read path.

Verifying an answer against source files catches what the model said. The
earlier problem is what the extractor is allowed to put into the store in the
first place. Every extracted record passes deterministic gates before it can
become structure, and the gates are stated as rejections rather than repairs: a
record that cannot be made into a well-formed assertion is dropped, not guessed
at. No placeholder structure, no vague predicates, no pronominal or conjunctive
entity names, no bare quantities as entities, and no inverting a predicate to
express a denial.

The placeholder rule is the one I would argue hardest for. A synthesised
stand-in written so the pipeline can continue is indistinguishable from genuine
extracted structure once stored, so a transient provider outage mid-run
permanently poisons the store; and nothing downstream can tell. Failing the
chunk loudly is the only option that keeps the artifact trustworthy. A wiki that
compounds is a wiki whose early errors compound too.


Two operations in the gist get easier once the structure is a graph, and
both for the same reason: they become computed rather than adjudicated.

The index. @kriss-b's observation that the Statement of Applicability
naturally becomes the index is, I think, the general case; every wiki grows
one, and maintaining it by hand is where the maintenance burden the gist warns
about actually lands. A clustered graph generates it instead. We cluster
entities into communities, summarise each, and then build a topic tree above
them by recursive supergraph clustering: one node per cluster, edge weights
summed across the cut, re-cluster, repeat per level.

The recursion is the part that matters, and it is worth being precise about
why. A hierarchy built by sweeping a resolution parameter gives you levels but
guarantees nothing about containment - children can sit outside their parents.
Any drill-down built on such a tree is lying: you click a topic and get
material that is not in it, or miss material that is. Recursing over the
supergraph guarantees nesting by construction, because level n+1 is built
from level n's own nodes. Parent summaries are synthesised from child
summaries rather than from raw relations, which caps the added LLM cost at
roughly the cluster count per level rather than re-reading the corpus.

The practical effect is a wiki index at several zoom levels that nobody
maintains, and that cannot drift from the pages beneath it, because it is
derived from them.

Lint's other half. The gist's lint pass looks for contradictions and gaps.
Contradictions are handled at write time, above. Gaps are the harder question,
because a wiki cannot normally tell you what it does not cover - absence has no
page. Recording which entities and communities each query touched turns that
into arithmetic: which regions of the corpus have received least retrieval
attention, and which entities have never been retrieved at all. Not an LLM
judging completeness, just a count over what was actually read. Storing a hash
of the query rather than its text keeps the telemetry from becoming a second
sensitive dataset.

The pattern across all three (conflicts, index, gaps) is the same, and
@wy-cats found the same edge from a different direction in discovering that
"pending ingest" works better as computed state than as an explicit flag. A
flag has to be maintained and can be wrong; a computed value cannot drift from
what it describes. An LLM sweep is the expensive way to rediscover something
the write path, a clustering, or a counter already knew.


One thing worth instrumenting, because it is invisible to every metric this
architecture naturally produces.

Between the retriever and the model sits an assembler that fits retrieved
material into a token budget. Any assembler that fills greedily and stops at
the first oversized passage will silently contribute nothing from that channel
and a wiki page or a filing section is routinely larger than a per-channel
budget. The model then answers "no information available" while holding the
answer, unread, in material that was retrieved successfully.

Retrieval metrics cannot see this. They are computed over what was found, not
over what survived into the prompt; recall@k reads as perfect throughout. On
LongMemEval the gap between clipping oversized passages and dropping them is
eight points of end-to-end accuracy - larger than most differences between
retrieval strategies, and attributable entirely to the step after retrieval.

Log how many retrieved characters actually reach the model. One number, and it
makes the whole failure class visible.


On markdown versus structure: markdown is readable, diffable and hand-editable,
and a graph gives all three up. Which side wins depends on whether your wiki is
read mostly by people or mostly by systems.

Both layers are open source and run entirely on PostgreSQL with pgvector for
similarity, ordinary tables for the graph, one backup, one transaction spanning
your graph and your application rows. No additional services.

Entity resolution across documents and structural negation so "not a
subsidiary of" is never retrieved as "subsidiary of" are the other two write-
path problems worth comparing notes on.

@kaimys

kaimys commented Sep 21, 2026

Copy link
Copy Markdown

We used the LLM Wiki to develop a mobile app, and it worked great. At first, we mainly used MCP to sync tickets into the wiki. Later, we stopped using the ticket system and relied entirely on the wiki. We added automated meeting transcripts generated by Google Gemini, then used an AI agent to turn them into meeting notes. The agent could reference the relevant tickets and even update them based on the transcript. It also drafted meeting agendas by collecting unanswered questions from open tickets for the next release and preparing them for the next product management meeting.

Later, we added automated ADRs (Architecture Decision Records) and release notes. The more we added, the more useful the wiki became. We even built a Kanban board and a release board, then open-sourced them as an Obsidian plugin. If you would like to try it, you can find it here: https://community.obsidian.md/plugins/dispatch

@orcosto-lab

Copy link
Copy Markdown

@mikhashev raised this in April as the open problem: how does the agent know to look for something it forgot it has? We run a small wiki (~15 pages, Claude and Codex writing to it, Civil 3D automation over COM plus a few home-built MCP servers) and ended up with two partial answers. Neither needs retrieval infrastructure.

1. Measure whether the agent consults the wiki, not only what it retrieves. Most numbers in this thread are retrieval quality. We audited 9 real working sessions instead, and found that the wiki was being read at session start but not mid-task. In 2 of those sessions the agent re-diagnosed from scratch a failure that already had a written lesson. Retrieval was fine; the lookup never happened. So we added an explicit trigger to the schema: when an external tool fails with a recognizable error signature and the first hypothesis didn't fix it, check the wiki before trying a second one. It's a rule rather than a hook, so it can still be skipped, but it gives a concrete moment instead of "consult when relevant".

2. Put the lesson where the agent can't avoid reading it: the tool description. For MCP servers we own, lessons that prevent a concrete failure go into the tool's description with a LESSON: prefix, in one line. The full context stays in the wiki page. The description is loaded every session and read right before the call, so it removes the "didn't know I should search" step entirely for that tool. It's the same idea as keeping docs as docstrings in the code (@tonydzi, above), applied to the one piece of text an agent always reads.

Limits: (2) only works for tools you control, and each line costs tokens in every session, so it's reserved for lessons that have already caused a real failure. (1) is a behavioural rule, and we still don't have a second audit showing how often it fires.

@giodra96

Copy link
Copy Markdown

Huge fan of this pattern!

I built an extension of LLM Wiki pattern specifically designed for software codebases: Project Wiki.

Feedback and contributions welcome: https://github.com/giodra96/project-wiki

@learnopengles

learnopengles commented Sep 21, 2026 •

Copy link
Copy Markdown

Here's my own personal mind map extension: LLM Wiki: Mindmap Extension.

Paste this into your LLM after the LLM Wiki pattern. It extends the wiki pattern into an "interactive map" of your mind, detailing your values, lessons, the things you care about, and what you hope to pass on.

I built this after seeing a tweet by Jen Zhu (the "tractable form of brain upload" as per Karpathy), along with inspiration from the idea of "digital gardens" (full sources in the gist).

@equationalapplications

Copy link
Copy Markdown

Jev and OpenJev can replace LLMs for classifying facts and relationships in the LLM Wiki pattern faster and more cheaply! I like to auto-classify my LLM Wiki into a Knowledge Graph using a fixed schema (ontology) for the types (nodes and edges). OpenJev is not an LLM, is is a general-purpose classifier, and it can handle this schematizing workload.

https://huggingface.co/openjev/openjev

@XBlueSky

Copy link
Copy Markdown

Follow-up to my July comment: after dogfooding Cortexes on real Claude Code sessions for a few more months, I increasingly think the hard part of an LLM wiki is not retrieval — it is making the compiled memory trustworthy.

A few failure modes surprised me:

  1. Capture pipelines should fail open. A transcript-filter payload drift silently produced unusable records for two weeks before I noticed it. The mistake was architectural: every compression/filtering stage is optional, but one exception could discard the whole session. The recorder now degrades per stage, preserves the raw content on failure, logs redacted diagnostics, and only moves a completed temp file into the vault after rendering succeeds.

  2. Provenance and identity are harder than summaries. SessionEnd can fire repeatedly on a growing transcript, so I reclaim strict-prefix snapshots — but only within the same repo. A real vault exposed a cross-repo prefix pair where deleting the shorter copy would have erased that repo's only record for months. Separately, a whitespace difference around a distillation marker changed source hashes and made already-distilled records look modified. I ended up validating the fix against 1,174 real Raws and their pre-marker git blobs.

  3. Ranking, score, and retrieval provenance are different things. Cortexes uses BM25 + vector search with RRF. Originally a BM25-only hit could rank #1 while exposing score: 0.0; the agent interpreted that as low confidence even though the ranking was correct. Fixing the prompt wasn't enough — hybrid-only hits now get their cosine backfilled, and repo/type/category filters are evaluated consistently by vector, BM25, and graph streams.

The pattern I keep seeing is that automatic capture is easy. The real system is everything that prevents silent loss, duplicate memory, stale state, misleading confidence, and accidental over-retrieval.

Current implementation, tests, and eval harness:
https://github.com/XBlueSky/cortexes

@mohitagw15856

Copy link
Copy Markdown

Built a version that puts a capture layer in front of the wiki: a Telegram bot you talk to in plain language ("ate paneer bhurji and 2 rotis", "mandarin for chopsticks", "spent 12 on lunch at Pret", "remind me at 6pm to call mum"). Claude classifies each message with structured outputs and writes into fixed-header markdown tables in the Obsidian vault (life/ for food, habits, money, words, reminders, people…). Claude Code stays the wiki maintainer exactly as in the gist: /ingest, /query, /lint, with CLAUDE.md as the contract. Same three layers, plus inbox/ and life/ in front, so the compounding starts from your day and not only from articles.

Things that surprised me:

  • Fixed table headers are the whole trick. Once the format is stable, the bot, Obsidian's Dataview and Claude Code all trust the same file.
  • One small typed call per capture (intent first, then the specific parser) beats prompting for prose.
  • Let the bot finish structured captures itself and keep the IDE session for synthesis. If everything waited for "process my inbox", nothing would get logged.
  • Timestamp by send time, not processing time, so a backlog after a closed laptop files under the right day.
  • A pinned "Today" message that edits itself after every capture did more for daily use than any feature.

Template with the bot, schema, Claude Code commands and an empty vault (no personal data): https://github.com/mohitagw15856/second-brain-template

@Roylin1003

Copy link
Copy Markdown

Thank you, Karpathy. Your LLM wiki led me here: https://roynexus.com/

Here's my approach:
https://knowledgetwin.substack.com/p/how-to-build-a-knowledge-twin

@oefi

oefi commented Sep 23, 2026

Copy link
Copy Markdown

So did anyone ran a host of benchmarks to verify the claims "Why this works"?

@plantdaddy99

plantdaddy99 commented Sep 23, 2026 via email

Copy link
Copy Markdown

@ojuschugh1

Copy link
Copy Markdown

https://github.com/ojuschugh1/sqz

  ███████╗ ██████╗ ███████╗
  ██╔════╝██╔═══██╗╚══███╔╝
  ███████╗██║   ██║  ███╔╝
  ╚════██║██║▄▄ ██║ ███╔╝
  ███████║╚██████╔╝███████╗
  ╚══════╝ ╚══▀▀═╝ ╚══════╝
  

Pre-injection context compression and session deduplication for AI coding agents

Real session stats: 3,003 compressions · 178,442 tokens saved · 24.7% avg reduction · up to 92% on output the model already has

Featured Featured on Hysen Labs

Crates.io npm PyPI VS Code Firefox JetBrains Discord Homebrew

Docs · Install · How It Works · Supported Tools · MCP Proxy · Benchmark · Changelog · Discord


AI coding agents spend most of their input tokens on tool output: build logs, test runs, git status, the same file read again after every edit. All of it goes into the context window raw, and it is re-sent on every turn after that.

sqz compresses tool output before it enters the context. Per-command formatters keep what an agent acts on (the failing test, the assertion, the file:line) and drop what it does not. Anything the model already has in context comes back as a short reference instead of the content. Source code, stack traces and secrets pass through untouched.

cargo install sqz-cli sqz-mcp     # or: brew install ojuschugh1/sqz/sqz · npm i -g sqz-cli · pipx install sqz
sqz init                          # hooks + MCP server + agent guidance for every client it finds

sqz: a 200-line file is served in full, a re-read of lines 41-80 becomes a line-range reference, cargo test output collapses to the failure, and sqz stats shows both

Real output from the release binary (assets/demo.tape). Regenerate with vhs assets/demo.tape.

Without sqz:                              With sqz:

cat auth.py (500 lines):   4,000 tokens    4,000 tokens  (source is served in full)
cargo test (1 failure):    1,200 tokens      120 tokens  (failure, assertion, file:line)
sed -n '40,80p' auth.py:     330 tokens       18 tokens  (§ref:…:L40-80§ to the read above)
git status, unchanged:       160 tokens       13 tokens  (§ref:…§)
─────────────────────────────────────      ─────────────
Total:                     5,690 tokens    4,151 tokens  (27% saved, nothing dropped)

Single Rust binary, deterministic, no LLM calls, works offline. Every compressed result can be recovered byte-exact with sqz expand.

Note

Name disambiguation: this repo, ojuschugh1/sqz, is an independent project and is not affiliated with any other similarly named or working tool or other compression projects that shorten "squeeze". If you installed sqz / sqz-cli / sqz-mcp from crates.io, npm, PyPI, or Homebrew, it comes from this repository.

Token Savings

24.7% average reduction across 3,003 real compressions ·
13-token refs for output the model already has ·
100% of agent-critical facts kept in the quality benchmark ·
0% by construction on source code, stack traces and secrets

One developer's week, measured from actual sqz gain output:

$ sqz gain
sqz token savings (last 7 days)
──────────────────────────────────────────────────
  04-13 │                              │   2,329 saved
  04-14 │                              │       0 saved
  04-15 │███                           │  12,954 saved
  04-16 │██                            │   9,223 saved
  04-17 │████                          │  14,752 saved
  04-18 │██████████████████████████████│ 105,569 saved
  04-19 │████████                      │  30,882 saved
  04-20 │█                             │   4,334 saved
──────────────────────────────────────────────────
  Total: 3,003 compressions, 178,442 tokens saved (24.7% avg reduction)

Per-command compression

Single-command compression (measured via cargo test -p sqz-engine benchmarks):

Content Before After Saved
Repeated log lines 148 62 58%
Large JSON array 259 142 45%
JSON API response 64 53 17%
Git diff 61 54 12%
Prose/docs 124 121 2%
Stack trace (safe mode) 82 82 0%

Session-level with dedup

Anything the model already has in its context comes back as a reference. Measured with the release binary against a throwaway database, using the fixtures in sqz/tests/quality_bench.rs and demo/:

Repeat Without sqz With sqz Saved
Unchanged git status run again 155 13 92%
Lines 41-80 of a 200-line file read earlier 305 16 95%
Same cargo test output, nothing changed 303 13 96%

A word on what actually repeats. A user who counted 92 of their own Claude Code sessions (script) found identical whole-file re-reads in 1 of 542 reads, and line-range re-reads in 8.4% of them. So sqz does not lead with "the same file read five times": the repeats that happen in practice are unchanged command output, line ranges of a file already read, and files re-read after a small edit, and each has its own reference type (details). Sessions with noisy command output and repeated commands see the biggest wins.

Install

Prebuilt binaries (no compiler required — works on every platform):

# macOS / Linux
curl -fsSL https://raw.githubusercontent.com/ojuschugh1/sqz/main/install.sh | sh

# Windows (PowerShell)
irm https://raw.githubusercontent.com/ojuschugh1/sqz/main/install.ps1 | iex

# Any platform via npm
npm install -g sqz-cli

# macOS / Linux via Homebrew
brew tap ojuschugh1/sqz
brew install sqz

Build from source via Cargo:

cargo install sqz-cli sqz-mcp

sqz-cli provides the sqz binary; sqz-mcp provides the MCP server. sqz-engine is a library dependency — it compiles automatically and does not need to be installed separately.

Build from source (cargo install sqz-cli) works too, but needs a C toolchain:

  • Linux: build-essential (apt) or equivalent
  • macOS: Xcode Command Line Tools (xcode-select --install)
  • Windows: Visual Studio Build Tools with the "Desktop development with C++" workload. Without these, cargo install fails with linker link.exe not found. If you don't already have them, use the PowerShell or npm install above instead.

Then initialize:

sqz init --global     # hooks apply to every project on this machine
# or
sqz init              # hooks apply to just this project (.claude/settings.local.json)

--global writes to ~/.claude/settings.json (the user scope per the
Anthropic scope table),
so the sqz hook fires in every Claude Code session on this machine. This is
the common case on first install. Your existing permissions, env,
statusLine, and unrelated hooks in ~/.claude/settings.json are
preserved — sqz merges its entries rather than overwriting.

Plain sqz init (project scope) is useful when you want sqz active only
inside one repo.

After installing, sqz doctor checks the whole chain — binary, database,
shell hook, which clients are detected vs actually routed through sqz, and
whether anything was compressed recently — and prints the fix for each gap
it finds.

See what sqz would have saved, before installing any hook

If you already use Claude Code or Kiro, your transcripts are on disk. Replay
them through the sqz engine and get a counterfactual estimate on your own
sessions — computed locally, nothing uploaded:

$ sqz discover --replay
sqz discover — counterfactual replay (last 7 days)
────────────────────────────────────────────────────────

  Sessions replayed:   24
  Tool outputs:        43
  Tokens (original):   116,598
  Tokens (compressed):  92,560

  Estimated avoidable: 24,038 tokens (20.6%)

  Largest opportunities:

    source                tokens in    avoidable
    web_fetch                 58,102       12,288
    read                      37,085       10,486

It scans ~/.claude/projects and ~/.kiro/sessions/cli, or point it
anywhere with --transcripts PATH. Each session replays against a fresh
throwaway cache, so dedup references never pretend to span sessions, and the
report says "estimate" because that's what it is: these outputs were not
compressed at the time.

Only using one agent? Pass --only (or --skip) to limit which
configs are written:

sqz init --only opencode              # just OpenCode, nothing else
sqz init --only opencode,codex        # OpenCode and Codex
sqz init --skip cursor,windsurf       # everything except Cursor and Windsurf

Accepted names: claude, cursor, windsurf, cline, gemini,
kiro, opencode, codex. Aliases (claude-code, gemini-cli, roo,
kiro-cli) also work. --only and --skip can't be combined.

Manual installation (preserve comments in your config)

sqz init round-trips your config file through a JSON parser to merge
the sqz entry, which drops any comments in your opencode.jsonc (and
the analogous JSON-with-comments files other tools accept). If you've
commented your config carefully and want to keep them, install by hand
instead.

OpenCode — two steps:

  1. Drop the plugin file in place. sqz prints the generated TS to
    stdout so you don't have to hand-write the path-escaping logic:

    mkdir -p ~/.config/opencode/plugins
    sqz print-opencode-plugin > ~/.config/opencode/plugins/sqz.ts
  2. Add the MCP entry to your existing opencode.jsonc yourself.
    Append this block inside the top-level mcp object (create the
    mcp object if it doesn't exist):

    "sqz": {
      "type": "local",
      "command": ["sqz-mcp", "--transport", "stdio"],
      "enabled": true
    }

Comments in the rest of your file stay put. OpenCode auto-discovers
the plugin file; no plugin array entry needed (adding one causes
double-loading, see issue #10).

Other tools — Claude Code, Cursor, Windsurf, Cline, Gemini CLI,
and Codex use plain JSON configs without comment support, so the
automated path is non-destructive there. Use sqz init --only <tool>
for those.

That's it. Shell hooks installed, AI tool hooks configured.

How It Works

sqz system architecture

sqz installs a PreToolUse hook that intercepts bash commands before your AI tool runs them. The output gets compressed transparently — the AI tool never knows.

Claude → git status → [sqz hook rewrites] → compressed output (85% smaller)

What gets compressed:

  • Shell output — 45+ per-command formatters (git, cargo, npm/pnpm/yarn, pytest, ruff, go test, docker, kubectl, aws, terraform, gradle, xcodebuild, adb logcat, dotnet, gh, grep/rg, tree, curl, and more)
  • JSON — strips nulls, compact encoding, TOON format
  • Logs — collapses repeated lines
  • Test output — shows failures only (state-machine parsers for Rust, Go, Python, JS, JVM)

What doesn't get compressed:

  • Stack traces, error messages, secrets — routed to safe mode (0% compression)
  • Your prompts and the AI's responses — controlled by the AI tool, not sqz

Supported Tools

Tool Integration Setup
Claude Code PreToolUse hook (transparent) sqz init
Cursor PreToolUse hook (transparent) sqz init
Windsurf PreToolUse hook (transparent) sqz init
Cline PreToolUse hook (transparent) sqz init
Gemini CLI BeforeTool hook (transparent) sqz init
Kiro Steering + MCP server sqz init
OpenCode TypeScript plugin (transparent) sqz init
Codex CLI AGENTS.md guidance + MCP server sqz init
Zed AGENTS.md guidance + MCP server sqz init
Copilot CLI preToolUse hook (transparent) sqz init
Copilot coding agent (CI) Repo-level hook, self-bootstraps in the cloud sandbox sqz init --ci + commit
Any MCP server sqz-mcp proxy wraps it, compresses its tool results see below
VS Code Extension Install from Marketplace
JetBrains Plugin Install from Marketplace
Chrome Browser extension ChatGPT, Claude.ai, Gemini, Grok, Perplexity
Firefox Browser extension Same sites

Compress Any MCP Server

sqz-mcp proxy sits between your agent and any stdio MCP server and compresses
what flows back: tool results go through the full sqz pipeline (dedup refs on
repeats, safe-mode for stack traces and secrets, error results untouched), and
verbose tool descriptions get compacted so tools/list stops eating your
context window. Everything stays reversible — the proxy injects an sqz_expand
tool so the agent can recover any original byte-exact.

Wrap a server by prefixing its command in your MCP config:

{
  "mcpServers": {
    "github": {
      "command": "sqz-mcp",
      "args": ["proxy", "--", "npx", "-y", "@modelcontextprotocol/server-github"]
    }
  }
}

Flags: --lazy-tools shortens every tool description to one sentence and
injects an sqz_tool_help tool that serves the full original docs on demand
(some MCP servers spend 10-40k tokens on tools/list alone); --no-desc
keeps descriptions verbatim; --no-cache disables dedup refs. Works with
every MCP client (Claude Code, Cursor, Windsurf, Zed, Codex, Kiro, ...)
because the client just sees a normal MCP server. Full guide, including the
server's own sqz_read_file / sqz_grep / sqz_list_dir tools and per-client
config: MCP context compression.

CLI

sqz init --global             # Install hooks for every project on this machine
sqz init                      # Install hooks for just this project
sqz init --only kiro          # Only configure Kiro (skip the rest)
sqz init --only opencode      # Only configure OpenCode (skip the rest)
sqz init --skip cursor        # Configure every agent except Cursor
sqz init --ci                 # Also write .github/hooks/sqz.json for Copilot agent CI runs
sqz compress <text>           # Compress (or pipe from stdin)
sqz compress --no-cache       # Compress without dedup (always full output)
sqz expand <ref>              # Recover original content from a §ref:HASH§ token
sqz compact                   # Evict stale context to free tokens
sqz reset                     # Clear dedup cache or compression stats
sqz gain                      # Show daily token savings (bar chart)
sqz gain --project .          # Per-project daily gains
sqz gain --days 30            # Last 30 days
sqz stats                     # Cumulative compression report (includes regret signals)
sqz stats --cost              # Estimated $ saved under a prompt-cached billing model
sqz stats --breakdown         # Per-command token usage breakdown
sqz stats --project .         # Stats for current project only
sqz stats --project list      # List all tracked projects
sqz stats --share             # Five-line share card sized for a paste into an issue or post
sqz doctor                    # Verify sqz is installed, wired into your clients, and active
sqz discover                  # Find missed savings (incl. your weakest-compressing commands)
sqz discover --replay         # Counterfactual replay of your agent transcripts (Claude Code, Kiro)
sqz recall "auth timeout"     # Full-text search everything sqz has compressed
sqz resume                    # Re-inject session context after compaction
sqz vizit                     # Live terminal dashboard (like htop for AI agents)
sqz hook claude               # Process a PreToolUse hook (Claude Code)
sqz hook kiro                 # Legacy; Kiro now uses steering + MCP (sqz init)
sqz print-opencode-plugin     # Print OpenCode plugin TS for manual install
sqz proxy --port 8080         # API proxy (compresses full request payloads)

Dedup Escape Hatch

When sqz sees the same content twice, it returns a compact §ref:HASH§ token
instead of the full text. Most models handle this fine, but some (e.g., GLM 5.1)
can't parse the ref format and loop. Four ways to work around this:

# 1. Recover original content from a ref
sqz expand a1b2c3d4              # prefix match
sqz expand '§ref:a1b2c3d4§'     # paste the whole token

# 2. Compress without dedup (per-invocation)
echo "..." | sqz compress --no-cache

# 3. Disable dedup globally (env var)
export SQZ_NO_DEDUP=1

# 3b. Or just shorten the ref freshness window (seconds, default 1800).
# Useful on clients without a compaction hook; 0 disables refs.
export SQZ_REF_TTL_SECS=300

# 4. MCP passthrough tool (returns input byte-exact, zero transforms)
# Available via tools/list when sqz-mcp is running

Recall: search everything sqz has seen

Everything that flows through sqz is indexed locally (SQLite FTS5, on your
machine, nothing leaves it). When your agent compacts its context and loses
that error message from an hour ago, search for it instead of re-running
the command:

$ sqz recall "connection refused"
  1. [git · 2h ago]  … curl: (7) Failed to connect: >>connection refused<< on port 5432 …
     full content: sqz expand 3f9a1c05e2d84b17

Agents get the same thing as the sqz_recall MCP tool.

Track Your Own Savings

Run sqz gain in your shell any time to see your own daily breakdown (see the
Token Savings section above for what the output looks like), and sqz stats
for the full cumulative report:

$ sqz stats
  📊 sqz compression stats
  ──────────────────────────────────────────────────

  178,442  tokens saved
  ↓  24.7% average reduction

  Compressions           3,003
  Tokens in              721,840
  Tokens out             543,398
  Tokens saved           178,442
  Avg reduction          24.7%

  🗄️  Cache
  ──────────────────────────────────────────────────
  Entries                43
  Size                   39.1 KB

Add --breakdown to see exactly which commands consume the most tokens:

$ sqz stats --breakdown

  🔍 Top Token Consumers
  ──────────────────────────────────────────────────────────────────────
  command               calls  tokens in        out    saved
  ──────────────────────────────────────────────────────────────────────
  dedup                   249      45541       3237      93%
  stdin                    51      30851      24289      21%
  auto                    132      18288       7740      58%
  echo                     17       1050        558      47%
  ls -la                    8        948        948       0%
  cargo build               7        170        145      15%
  git status                4         56          8      86%
  ──────────────────────────────────────────────────────────────────────

Honest accounting: cost and regret

Token reduction is not automatically billed-cost reduction — with provider
prompt caching, most context re-transmits at a ~90% discount, and compression
that drops something the agent needed costs extra turns instead
(arXiv:2607.12161). sqz measures both
sides instead of hand-waving:

  • sqz stats --cost estimates dollars saved under a prompt-cached billing
    model (cache write ×1.25 once, cache read ×0.10 per subsequent turn), with
    every assumption printed and overridable (--price-in, --reread-turns,
    --cache-write-mult, --cache-read-mult). This is the conservative
    estimate: without caching the same savings would bill at full input price
    on every turn.
  • sqz stats reports regret signals: quick re-runs (the agent re-produced
    byte-identical output within 2 minutes — the repeat bought nothing) and
    ref expands (§ref§ tokens recovered to original bytes). Both are proxies
    for "compression dropped something the model needed." If one command
    keeps showing up, its formatter needs work — file an issue.

sqz's design already avoids the failure modes that study measured:
compression is deterministic and query-agnostic (never invalidates provider
prompt caches), only new tool output is touched (history is never rewritten),
piped or redirected commands are never intercepted, and the 16-token net-win
gate skips compressions that wouldn't pay for their own markers.
image

Per-project filtering:

sqz stats --project .           # stats for current project only
sqz stats --project list        # list all tracked projects
sqz gain --project .            # daily gains for current project
sqz gain --days 30              # last 30 days instead of 7
sqz gain --days 30 --project .  # combine both

Stats are stored locally in SQLite under ~/.sqz/sessions.db — nothing leaves your machine.

How Compression Works

  1. Per-command formatters — 45+ commands across 12 ecosystems get purpose-built compression:

    Ecosystem Commands
    Git status, log, diff, show, stash, remote, fetch, push, pull, commit
    Rust cargo build/test/clippy/check/nextest
    JavaScript npm/pnpm/yarn/bun install/test/audit/outdated, tsc, eslint, vitest
    Python pytest, ruff, mypy, pip
    Go go test (incl. -json stream), go build, go vet, golangci-lint
    Cloud aws, terraform plan/apply/init, gcloud
    Containers docker/podman ps/images/build, kubectl get/describe/logs/apply
    JVM gradle build/test, maven
    Mobile xcodebuild (build + test), adb logcat
    .NET dotnet build/publish/test
    System grep/rg, tree, find/fd, ls, curl/wget
    GitHub gh pr/issue/run (JSON + table)

    Unknown commands fall through to the generic compression pipeline — no output is ever left uncompressed.

  2. Table compactor — aligned-column output from tools without a dedicated formatter (ps aux, netstat, database CLIs) collapses its padding runs to two-space separators. The detector is strict — indented lines, code, YAML, and JSON never match.

  3. Structural summaries — code files compressed to imports + function signatures + call graph (~70% reduction). The model sees the architecture, not implementation noise.

  4. Dedup cache — SHA-256 content hash, persistent across sessions. Byte-identical output the model already has = 13-token reference. A line range of a file the model already has in full (sed -n '40,80p', sqz_read_file with offset/limit) = §ref:HASH:L40-80§. A re-read after a small edit = only the changed lines. References are only served while the original is still in context (30-minute window, tunable via SQZ_REF_TTL_SECS, reset on compaction).

  5. JSON pipeline — strip nulls → project out debug fields → flatten → collapse arrays → TOON encoding (lossless compact format)

  6. Safe mode — stack traces, secrets, migrations detected by entropy analysis and routed through with 0% compression

Measured quality, including what each stage drops and what it never touches: quality benchmark. For the full technical details, see docs/.

Configuration

# ~/.sqz/presets/default.toml
[preset]
name = "default"
version = "1.0"

[compression.condense]
enabled = true
max_repeated_lines = 3

[compression.strip_nulls]
enabled = true

[budget]
warning_threshold = 0.70
default_window_size = 200000

Per-project database

Everything sqz persists (stats, dedup cache, sessions) lives in one SQLite
file, ~/.sqz/sessions.db by default. Set SQZ_DB_PATH to keep it
per-project instead:

# e.g. in the project's .envrc (direnv)
export SQZ_DB_PATH="$PWD/.sqz/sessions.db"

Every surface honors it — the shell hook, sqz stats/gain/expand, and
the MCP server (set it under env in your MCP config). Prefer absolute
paths: relative values resolve against whatever directory the process runs
from. The parent directory is created if missing (mode 0700, like ~/.sqz).

Privacy

  • Zero telemetry — no data transmitted, no crash reports
  • Fully offline — works in air-gapped environments
  • All processing local

Development

git clone https://github.com/ojuschugh1/sqz.git
cd sqz
cargo test --workspace
cargo build --release

License

Elastic License 2.0 (ELv2) — use, fork, modify freely. Two restrictions: no competing hosted service, no removing license notices.

Links

Star History

Star History Chart

Contributors

Thanks to everyone who has contributed code, fixes, and ideas to sqz:

sqz contributors

And to everyone who filed the detailed bug reports behind our fixes. Precise repros make this project better with every release. Want to join them? PRs are reviewed fast: see the open issues to get started.

Contributor grid made with contrib.rocks.

@equationalapplications

equationalapplications commented Sep 23, 2026 •

Copy link
Copy Markdown

Take advantage of the speed and low price of a general purpose, system one type classifier like Jev or OpenJev to classify the LLM Wiki node types in a knowledge graph using Curated Thoughts.
https://github.com/equationalapplications/curated-thoughts

Screenshot 2026-09-23 at 3 55 09 PM

@plantdaddy99

plantdaddy99 commented Sep 24, 2026 via email

Copy link
Copy Markdown

@jouskaee-tech

Copy link
Copy Markdown

useful!!

@caspianmd

Copy link
Copy Markdown

This has been a huge inspiration! I built a second brain LLM wiki with UI in Caspian that you can clone for free :)

https://second-brain-llm-wiki.sites.caspian.md

HTYaJ06bsAEs9mf

@Sistema2D

Copy link
Copy Markdown

FrameCode-VibeWork v0.20.0 is now available

Capa FrameCode com subtítulo em inglês

The new version of FrameCode-VibeWork simplifies the framework structure and improves upgrade integrity, validation, and rule triggering.

Key highlights:

  • Leaner package: from 273 to 103 files, with an approximately 30% reduction in installed size.
  • A single PROJECT.md replaces seven profiles, consolidating policies, templates, and skills.
  • Upgrades that protect local changes and provide controlled removal of obsolete files.
  • More explicit criteria for triggering rules and skills, plus validation of architectural decision records (ADRs).

Available in Portuguese, English, Spanish, and German. Upgrading requires migration according to the release instructions.

Explore the changes and download v0.20.0.

@ShootJackal

Copy link
Copy Markdown

Obelyth Cortex v2.0.2: every cited quote checked against a Git commit

The hard part of an LLM-maintained wiki is trusting what it tells you later. Cortex runs this pattern as one MCP server over a private Git repo of Markdown, and a check with no model in it confirms each quote is in the cited note at that commit.

01-hero
The gist Obelyth Cortex
Wiki, index, log Markdown in a private GitHub repo. Every note write is a commit; INDEX.md rebuilds; captures land in a daily log.
Ingest Edit in place, append to the log, or import a folder. No raw-source compiler yet.
Query brain_ask answers, then the quote is checked. Quotes from retracted passages come back SUPERSEDED.
Lint Flags stale stamps, superseded links, secret-shaped text.
Schema One MCP server for every client; setup wires Claude Code, Cursor, Gemini CLI, claude.ai and ChatGPT.
Trust Guests can only ask and propose. Nothing they send is committed until you accept it.

One we got wrong: a note could once forge a file boundary and earn VERIFIED for a fabricated answer. The reader now cites an opaque per-commit tag, never a path (lib/ask.ts).

@crajah, on the write path: we refuse rather than repair where we can. An oversized capture is refused, never truncated; a capture returns its commit SHA, and the model is told never to claim a save without one. Trusted writes don't pass extraction gates like yours yet.

Check it yourself. Passing tests: 439 at v1.0.0 (Aug 11) → 3,068 at v2.0.2 (Oct 1). The 139 database tests run separately against real Postgres and Valkey, and a single skip fails the release. Each release re-runs its checks at the tag, waits for a maintainer, and ships an attested tarball:

gh release download v2.0.2 -R Obelyth/cortex -p '*.tar.gz'
gh attestation verify cortex-v2.0.2.tar.gz -R Obelyth/cortex

Self-hosted, AGPL-3.0, starts blank. Import a copy of your notes and tell us where a stamp is wrong.

Repo · v2.0.2 release · Roadmap

@gimalay

gimalay commented Oct 4, 2026

Copy link
Copy Markdown

Maybe I'm missing something, so please explain if I am wrong: How does this differ from teaching your agent to use iwe-org/iwe ?

How to use this pattern with iwe: https://gist.github.com/gimalay/6415b269dd9188060c4ea60ede5eb7bb

@alfabarisl18-ops

Copy link
Copy Markdown

Yes—if you teach the agent the same workflow, IWE can implement this pattern. Karpathy’s gist describes the methodology: preserve raw sources, incrementally maintain a derived wiki, and use ingest/query/lint rather than repeatedly synthesizing from raw documents. IWE supplies concrete tools for working with the Markdown graph; connecting it alone doesn’t specify that workflow.

Your linked example makes the connection explicit: it retains the three layers and operations, while adding inclusion-link navigation, deterministic schema/orphan checks, and renames that update references. I’d therefore describe it as an implementation of the pattern. As your example notes, those checks validate structure; contradictions, stale claims, and factual accuracy still need review.

@bg-numarqe

bg-numarqe commented Oct 5, 2026 •

Copy link
Copy Markdown

Karpathy's got the right idea — knowledge should compound instead of getting re-derived every session. He's got the wrong tool for coding and developing though.

Markdown's fine for generalist knowledge. Code isn't generalist knowledge though, is it. It's relational. A folder of .md files falls apart the minute you need to know "what calls this function" across fifty files. And it goes stale the second someone renames a thing. At any real scale you've just built a bigger pile of stuff the LLM has to grep through. Bit of a dog's dinner.

I pointed Synapse MCP at my codebase — the one that does this with a live graph instead of markdown. It just indexed the lot. Didn't have to think about it. The LLM stopped faffing about trying to find where things live and just got on with fixing stuff. That's about all there is to it really.

Anyway @spasm-myelixlabs — any chance of a discount code for the pro version? Asking for a mate. The mate's me.

@spasm-myelixlabs

Copy link
Copy Markdown

Karpathy's got the right idea — knowledge should compound instead of getting re-derived every session. He's got the wrong tool for coding and developing though.

Markdown's fine for generalist knowledge. Code isn't generalist knowledge though, is it. It's relational. A folder of .md files falls apart the minute you need to know "what calls this function" across fifty files. And it goes stale the second someone renames a thing. At any real scale you've just built a bigger pile of stuff the LLM has to grep through. Bit of a dog's dinner.

I pointed Synapse MCP at my codebase — the one that does this with a live graph instead of markdown. It just indexed the lot. Didn't have to think about it. The LLM stopped faffing about trying to find where things live and just got on with fixing stuff. That's about all there is to it really.

Anyway @spasm-myelixlabs — any chance of a discount code for the pro version? Asking for a mate. The mate's me.

@bg-numarqe

Haha thanks for giving it a spin, for your 'mate' and any of our friends here. Use KARPATHY100 for the pro version. FREE for 6 months. My only ask is to give us your feedback.

Synapse is FREE - it can do everything - for code and your repos laid out in the gist and more, but without markdown.

Pro adds write safety (stops the LLM doing something stupid before it hits disk), test intelligence, and memories - if your coding full time, these are useful to you.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment