Agents Need Documents They Can Check
Kevin Liao's essay Agents Don't Need Memory. They Need Documentation. argues that memory plugins for coding agents are just retrieval: snippets cut from old transcripts, picked by similarity and pushed into every prompt. His alternative, operator-memory, is a set of Markdown documents the agent reads before a task and updates after it. I agree with most of his diagnosis. It describes a memory system I work on. What his answer leaves out is a way to know whether a document is still true.
The agent reads a document. Every line in it can show where it came from and whether it is still true.
What he gets right
Liao lists five problems with memory plugins:
- Similarity is not correctness. The closest memory may be wrong or out of date.
- Snippets lose the context they came from.
- The past is treated as truth, while the code keeps changing.
- An agent cannot search for something it does not know exists.
- Nobody can audit thousands of embeddings by reading them.
These are real. Astrocyte, the open-source memory library I work on in my own time, has a Claude Code hook that does exactly what he describes: on every non-trivial prompt, it recalls memories by similarity and adds them to the prompt. His critique applies to it directly.
He is also right that documents are the better interface for coding agents. A person can read them, review them in a pull request and fix them by hand.
Where his evidence stops
His evidence is one developer's year of use on his own projects. There is no measurement. That does not make him wrong, but it leaves several cases open.
- Some memory has no author. In a codebase, someone has reason to keep the docs current. Conversational memory, something a user said in session 37 of 500, has no one to write it down and too much volume to read in full.
- Documentation is memory with a different write policy. The agent writes at the end of a task, while it still has the full picture. That is a good policy. It is also the same job a background memory process does, run in the foreground instead.
- Documents fail in their own ways. Agent-written docs drift. When a model rewrites a page, it can drop information and leave no record of what it dropped. Two agents editing one file at the same time conflict.
- Past a certain size, reading the index is retrieval again. Once the documents outgrow the context window, "consult the brain" means the model choosing which pages to open. That is search, done by the model.
What we found in our own system
Liao's critique made me check Astrocyte against each of his points. Some of what we found agrees with him more than I would like.
They are internal self-evaluations on the LongMemEval-S benchmark, answered and judged by Claude Haiku. They are not scores from a public leaderboard, and they cannot be compared with leaderboard or vendor numbers.
Similar is not current. A step that reorders recall results was rebuilding each result without its event date. So every recalled memory carried the date it was stored, not the date it happened, and nothing looked wrong. After the fix, recalled results carried 12,500 real event dates across 426 distinct days, up from zero. The number falling back to the storage date dropped from 1,500 to zero.
The store was hard to audit. In one benchmark run on 4 October, items that had been retried after a restart held 22.5% duplicate rows, against 3.8% for clean items. The worst case was 43%. Only a SQL query showed it. Retries were storing half-finished batches a second time, and duplicate detection forgot what it had seen whenever the process restarted. Both are fixed in v0.18.0. The duplicates did not change that run's result. We found them because we looked.
Similarity cannot see "not". "The user is allergic to peanuts" and "The user is not allergic to peanuts" score 0.92 for similarity. That is close enough for a system to treat a correction as a repeat of the fact it corrects, and drop it. Duplicate detection now refuses a match when the two texts differ by a negating word.
Snippets and context. Astrocyte stores the original text of each chunk rather than short facts extracted by a model, to keep the surrounding context. That is a design choice. We have not yet measured it on its own.
His point
What we found in Astrocyte
Status
Similar is not current
Recalled memories showed when they were stored, not when they happened
Fixed
The store cannot be audited
Retried items held 22.5% duplicate rows, seen only through SQL
Fixed in v0.18.0
Similarity misses meaning
“Allergic” and “not allergic” scored 0.92 alike; a correction could be dropped
Fixed
Snippets lose context
We store original chunk text, not extracted facts
Not yet measured
Agents can’t search for unknowns
Our hook recalls by what the user typed
Open: table of contents proposed
The numbers themselves need checking. Our early runs used 30 questions. Re-judging the same answers at that size moved the score by over 3 points, with a margin of error of about 17 points either way. At 250 questions the score was 59.2–60.0%, with a 95% confidence interval of 53.8% to 65.9%. Even at 50 questions, two judging passes over identical answers differed by 4 points.
That is why self-reported results deserve caution, ours included. On the public Agent Memory Leaderboard, where every system runs under the same harness, MemOS scores 45.89. Its own reports put it at 88 to 89.
The corner neither one reaches
Memory and documentation make opposite trade-offs. Memory takes no effort, but nobody can see inside it. Documentation can be read, reviewed and handed over, but someone has to keep writing it. People trust memory by its results: it remembered my preference. They trust documents by inspection: I can see it says that. In a team, inspection matters more. You cannot review, approve or hand over something you cannot see.
More control
Best of both
Captured automatically
Reviewed like documents
Trusted by results and inspection
Documentation
High effort, high control
Trusted by inspection
Less control
Memory
Low effort, low control
Trusted by results
Worst of both
High effort, low control
Hidden and hard to keep
Less effort to keep current
More effort
- ↑ Memory moves up: show what it knows as readable pages.
- ← Documentation moves left: agents draft changes, people approve them.
The useful corner is the empty one: knowledge captured as easily as memory and checked as easily as a document. Liao's approach gets part of the way, because the agent does the writing. What it lacks is a way to check what the agent wrote. The proposal below aims for the rest.
A proposal: documents backed by evidence
We have written up a design that keeps Liao's best idea and adds what his approach drops. It is a proposal. None of it is built yet.
The agent reads documents, as Liao suggests. Each statement in a document is a claim, and each claim carries:
- Citations to the conversations or work that support it.
- Anchors to what it describes: a file, at a commit, down to the lines.
- Trust. Whether a person stated it, the agent checked it, or the agent inferred it.
- Status. Current, suspect, contradicted or replaced.
Claim
Sessions expire after 30 minutes of inactivity.
- Page
- auth.md › Sessions
- Citations
- 2 conversations, 12 Sep 2026
- Anchor
- src/auth/session.py @ a1b2c3d, lines 40–58
- Trust
- Checked by the agent against the code
Suspect: src/auth/session.py changed since a1b2c3d. Queued for re-checking.
The read path changes too:
- At session start, the agent gets a table of contents. It cannot ask about something it does not know exists, but it can see a map.
- When the agent opens a file, it gets the claims anchored to that file. The match is exact, not by similarity. It is triggered by what the agent is doing, not what it typed.
- Before a claim is shown, it is checked against the code. If the file has changed since the claim was anchored, the claim is marked "may be stale" and queued for re-checking.
- Similarity search stays, as a fallback for everything else.
At the end of a task, the agent proposes changes to the documents, with citations. That is Liao's write-after-work step, with a record of why each change was made.
Documentation loop
operator-memory
- Read the relevant documents
- Do the work
- Update the documents
↺ next task
Anchored documents
Astrocyte design, proposed
- Session start: read a table of contents
- Open a file: get the claims anchored to it
- Check each claim against the code; flag “may be stale”
- Do the work
- Propose document changes, with citations
↺ next task
How we would test his claim
His essay offers no measurement, so the test has to start there.
- Build his system as the baseline. Markdown documents only, no retrieval, the agent told to consult and update them.
- Test staleness directly. Write memories, change the code underneath them, and measure how often a stale memory misleads the agent. Nobody publishes this today.
- Use a shared benchmark. The Agent Memory Leaderboard's coding-memory track compares systems on the same tasks.
- Report cost and time with accuracy. Updating documents at the end of every task has a cost too.
Every comparison runs on identical questions, with repeated judging and a margin of error on every number.
What this means for Astrocyte
Liao's essay changes where Astrocyte is headed. It should sit behind documents as their evidence. Competing with them would be the wrong goal.
- The prompt hook stops being the default. Recalling by similarity on every prompt is the habit he describes. The plan is to replace it with a table of contents at the start of a session and the claims tied to whatever file the agent opens, keeping similarity as a fallback.
- Capture stays automatic. Conversations and work are kept as evidence without anyone writing them up. That is the part memory does well.
- What the agent knows becomes pages people can read. No one should have to query a database to see it.
- Agents draft, people approve, in the tools the team already uses. Proposed changes go out as pull requests on the repository's Markdown first, then as suggestions in documentation tools. Astrocyte does not need its own review screen.
- It should work with the documents you already have, including Markdown written by operator-memory, and add sources and staleness checks to them.
- It should draw on the search systems you already run, treating their results as evidence with sources, never as raw text for the agent to trust.
- Results come before claims. We will build Liao's system as the baseline and publish the comparison with its margins of error, whichever way it goes.
None of this is built yet. It comes after the benchmark work currently under way, and the design is public for anyone to read and question.
What to ask of any agent memory
Whether you build on documents, memory or both, three questions separate a system you can trust from one you hope is right:
- Where did this come from? Every claim the agent reads should point to its source.
- Is it still true? The system should know when the thing a claim describes has changed.
- How would you find out it was wrong? If the only way to audit it is a SQL query nobody runs, nobody will.
Liao is right that agents should read documents. They should also be able to check them.
