Blog
Karpathy's LLM wiki: how it works, and what a team needs to add
Published: October 6, 2026
Karpathy's LLM wiki is a folder of markdown pages that an LLM agent writes and keeps current from your sources, so knowledge builds up instead of being retrieved from scratch on every question. He published it as a gist on 2026-04-04, for personal knowledge bases. For one person it works. A team sharing one needs three things the gist leaves out.
Everything below about the pattern comes from the gist itself, which has one revision and has not been edited since. It is a different thing from the Karpathy CLAUDE.md, a set of coding guidelines that carries his name.
What is Karpathy's LLM wiki?
It is a pattern, which he calls an idea file: you paste it into Claude Code, Codex or another agent, and the agent builds the specifics with you. The contrast is with RAG, where the model rediscovers the same fragments on every query. Here the model files what it learns once, cross-links it, and keeps it current. In his words, “You never (or rarely) write the wiki yourself”. He keeps the agent open on one side and Obsidian on the other.
| Layer | What it is | Who writes it |
|---|---|---|
| Raw sources | Articles, papers, images and data. Never modified; the source of truth | You curate them |
| The wiki | Summary, entity, concept and comparison pages, cross-linked | The LLM, entirely |
| The schema | A CLAUDE.md or AGENTS.md describing the structure, conventions and workflows | You and the LLM, over time |
Three operations run it. Ingest: drop in a source and the agent summarizes it, updates the pages it touches (often 10 to 15) and logs it. Query: the agent answers from the wiki with citations, and a good answer is filed back as a new page. Lint: when you ask, it checks for contradictions, stale claims that newer sources superseded, orphan pages and missing links.
Two files hold it together. index.md lists every page with a one-line summary, and the agent reads it first on every query; he reports that this works at about 100 sources and hundreds of pages without embeddings, and suggests a local search engine such as qmd beyond that. log.md is an append-only record of ingests, queries and lint passes.
Which tools implement the LLM wiki?
Many, most within two weeks of the gist. The ones with the most use, as of 2026-10-06:
| Tool | What it is | Built for |
|---|---|---|
| Karpathy LLM Wiki (Obsidian plugin) | 58,173 downloads. Pages marked reviewed are protected from overwrite | One person; its README says it has no team collaboration |
| nashsu/llm_wiki | 20,226 stars. Desktop app with a review queue and contradiction checks at ingest | One person |
| claude-obsidian | 15,373 stars. Tracks contradiction, freshness and review state per page | One vault |
| llm-wiki-compiler | 2,162 stars. Approve or reject queue; holds pages on a contradiction | One person |
| synthadoc | 1,556 stars. Pages move from draft to active to stale or contradicted to archived | Small teams, 3 to 20 |
| codealmanac | 997 stars. A codebase wiki kept by coding agents, reviewed in git | Engineering teams |
Look at the third column. Nearly everything built on the gist kept its first word: personal.
Does an LLM wiki actually work?
For synthesis, there is early evidence that it does, at a cost. A preregistered comparison, arXiv:2605.18490 (v3, 2026-10-02), had an LLM-compiled wiki and single-round vector RAG answer the same 13 questions over 24 papers. The wiki scored much better at connecting findings across papers, spent about 21 times more tokens per query, and showed no citation-support advantage on human labels.
The index is the weak assumption. In arXiv:2607.04576 (2026-07-06), run on a real 709-page LLM-maintained wiki, a pilot found that a capable tool-using agent never loaded the index; it guessed a page's path from the question and read it directly. Both are single studies on small corpora, not yet peer reviewed.
On Hacker News, the field reports were good for one person (one reader turned three books into 210 pages, and the wiki caught a real contradiction between two of them). The objections were about time: a point beyond which the agent cannot keep up, and a lint that compares every page with every other.
What breaks when a team shares an LLM wiki?
The gist gives teams one bullet: an internal wiki fed by Slack threads, meeting transcripts and project documents, “possibly with humans in the loop reviewing updates”. It also notes that git gives you collaboration for free. Git records who changed a line. It does not say whether the line is true.
When teams tried it, the failures had a pattern. In the Show HN thread for a wiki that agents maintain, one commenter warned that six months in, entries are “confidently wrong” and lint cannot tell which; another, that bad entries get cited by other agents; a third put it as “Everyone is writing. Nobody is reading.” Taktile, which runs one over Notion and Google Drive, found its sources are edited by people who do not know a wiki depends on them, and that errors in an index the LLM writes itself compound.
The measured version is in work on shared agent memory, which is not about wikis but has the same shape: many writers, one store, many readers. In arXiv:2609.30813, once an uncontested false belief was in the shared store, agents asserted it in 97 to 99 percent of probes. Gating what got in by the declared source type held false adoption to 6 to 9 percent, against 22 to 47 percent for the other policies tested.
What does a team version of the LLM wiki need?
Three things, and each replaces a habit one person can get away with.
| One person | A team | Why |
|---|---|---|
| The LLM writes, you watch it happen | A gate: anyone drafts, a person or a policy you set promotes | An uncontested false claim in a shared store is repeated almost every time |
| Lint for stale claims when you remember | Dates on every claim, from its source, and retirement when a newer claim replaces it | Ten editors means the wiki goes stale faster than anyone lints it |
| An index the agent may or may not read, and a log of what was written | The relevant pages delivered with the task, and a count of which ones were read and used | Without a count, nobody knows which pages matter, or which to delete |
Take one decision through it. In #eng-payments the team agrees that webhook handlers acknowledge with a 200 and enqueue, and never do the work inline. In a personal wiki, it gets in if the one person who ingests that thread notices it. In a team wiki without a gate, it gets in along with every half-decision from the same thread, and agents cite them all alike. With a gate, a date and a count, it is approved, dated to the day it was said, handed to the session editing a webhook handler, retired when the queue is replaced, and you can see how often it was used.
One caution from a builder of these tools: a read count is not a truth signal. “Frequently asked is not the same as true”, as the karpathy-llm-wiki README puts it. A count tells you what is used, so you can retire what is not. It does not tell you what is right.
Where Harbor fits
Harbor is that team version for one kind of page: the decisions an engineering team makes. It reads where the team decides: Slack threads, the first source Karpathy lists for a team wiki, plus pull request reviews and docs. A person approves what it finds, or a policy you set approves the routine ones. A rule that contradicts a live one always waits for a person. Each rule is dated by the message it came from and can carry an end date, and when one replaces another, the sessions that used the old one are told. Claude Code, Codex and Cursor get the rules that apply to the task at session start, and every rule shows how often it was served and how often an answer cited it. A cite is evidence that a rule was used, not proof that it was obeyed; the method is in served vs cited.
For a research wiki of your own papers or books, use the gist as written. It is very good at that.
Questions
What is Karpathy's LLM wiki?
A pattern Andrej Karpathy published as a gist on 2026-04-04. An LLM agent builds and maintains a folder of linked markdown pages from your sources, with an index, a log, and a schema file such as CLAUDE.md or AGENTS.md, so knowledge builds up instead of being retrieved from scratch on every question.
How is an LLM wiki different from RAG?
RAG retrieves chunks of raw documents at query time and rebuilds the answer every time. An LLM wiki compiles the sources once into linked pages the model keeps current, and answers from those. One preregistered study found the wiki much better at connecting findings across papers, at about 21 times the tokens per query (arXiv:2605.18490).
Does an LLM wiki need a vector database?
Not at small scale. Karpathy reports that an index.md file listing every page works at about 100 sources and hundreds of pages. Beyond that he suggests a local search engine such as qmd.
Can a team use Karpathy's LLM wiki?
Yes, but the gist is written for one person. A team needs a gate on what goes in, dates and retirement for claims that go stale, and a count of which pages agents actually read, because one wrong entry in a shared store is repeated by every agent that reads it.
Is the LLM wiki the same as the Karpathy CLAUDE.md?
No. The LLM wiki is a knowledge-base pattern from Karpathy's own gist. The Karpathy CLAUDE.md is a set of four coding principles that a GitHub user distilled from one of his posts; it borrows his name.