Skip to main content

Agents Need Documents They Can Check

· 9 min read
Calvin Cheng
Shape what gets built and the value it creates.

Kevin Liao's essay Agents Don't Need Memory. They Need Documentation. argues that memory plugins for coding agents are just retrieval: snippets cut from old transcripts, picked by similarity and pushed into every prompt. His alternative, operator-memory, is a set of Markdown documents the agent reads before a task and updates after it. I agree with most of his diagnosis. It describes a memory system I work on. What his answer leaves out is a way to know whether a document is still true.

The agent reads a document. Every line in it can show where it came from and whether it is still true.

What he gets right​

Liao lists five problems with memory plugins:

  1. Similarity is not correctness. The closest memory may be wrong or out of date.
  2. Snippets lose the context they came from.
  3. The past is treated as truth, while the code keeps changing.
  4. An agent cannot search for something it does not know exists.
  5. Nobody can audit thousands of embeddings by reading them.

These are real. Astrocyte, the open-source memory library I work on in my own time, has a Claude Code hook that does exactly what he describes: on every non-trivial prompt, it recalls memories by similarity and adds them to the prompt. His critique applies to it directly.

He is also right that documents are the better interface for coding agents. A person can read them, review them in a pull request and fix them by hand.

Where his evidence stops​

His evidence is one developer's year of use on his own projects. There is no measurement. That does not make him wrong, but it leaves several cases open.

  • Some memory has no author. In a codebase, someone has reason to keep the docs current. Conversational memory, something a user said in session 37 of 500, has no one to write it down and too much volume to read in full.
  • Documentation is memory with a different write policy. The agent writes at the end of a task, while it still has the full picture. That is a good policy. It is also the same job a background memory process does, run in the foreground instead.
  • Documents fail in their own ways. Agent-written docs drift. When a model rewrites a page, it can drop information and leave no record of what it dropped. Two agents editing one file at the same time conflict.
  • Past a certain size, reading the index is retrieval again. Once the documents outgrow the context window, "consult the brain" means the model choosing which pages to open. That is search, done by the model.

What we found in our own system​

Liao's critique made me check Astrocyte against each of his points. Some of what we found agrees with him more than I would like.

What these numbers are

They are internal self-evaluations on the LongMemEval-S benchmark, answered and judged by Claude Haiku. They are not scores from a public leaderboard, and they cannot be compared with leaderboard or vendor numbers.

Similar is not current. A step that reorders recall results was rebuilding each result without its event date. So every recalled memory carried the date it was stored, not the date it happened, and nothing looked wrong. After the fix, recalled results carried 12,500 real event dates across 426 distinct days, up from zero. The number falling back to the storage date dropped from 1,500 to zero.

The store was hard to audit. In one benchmark run on 4 October, items that had been retried after a restart held 22.5% duplicate rows, against 3.8% for clean items. The worst case was 43%. Only a SQL query showed it. Retries were storing half-finished batches a second time, and duplicate detection forgot what it had seen whenever the process restarted. Both are fixed in v0.18.0. The duplicates did not change that run's result. We found them because we looked.

Similarity cannot see "not". "The user is allergic to peanuts" and "The user is not allergic to peanuts" score 0.92 for similarity. That is close enough for a system to treat a correction as a repeat of the fact it corrects, and drop it. Duplicate detection now refuses a match when the two texts differ by a negating word.

Snippets and context. Astrocyte stores the original text of each chunk rather than short facts extracted by a model, to keep the surrounding context. That is a design choice. We have not yet measured it on its own.

His point

What we found in Astrocyte

Status

Similar is not current

Recalled memories showed when they were stored, not when they happened

Fixed

The store cannot be audited

Retried items held 22.5% duplicate rows, seen only through SQL

Fixed in v0.18.0

Similarity misses meaning

“Allergic” and “not allergic” scored 0.92 alike; a correction could be dropped

Fixed

Snippets lose context

We store original chunk text, not extracted facts

Not yet measured

Agents can’t search for unknowns

Our hook recalls by what the user typed

Open: table of contents proposed

His critique, checked against our own system.

The numbers themselves need checking. Our early runs used 30 questions. Re-judging the same answers at that size moved the score by over 3 points, with a margin of error of about 17 points either way. At 250 questions the score was 59.2–60.0%, with a 95% confidence interval of 53.8% to 65.9%. Even at 50 questions, two judging passes over identical answers differed by 4 points.

That is why self-reported results deserve caution, ours included. On the public Agent Memory Leaderboard, where every system runs under the same harness, MemOS scores 45.89. Its own reports put it at 88 to 89.

The corner neither one reaches​

Memory and documentation make opposite trade-offs. Memory takes no effort, but nobody can see inside it. Documentation can be read, reviewed and handed over, but someone has to keep writing it. People trust memory by its results: it remembered my preference. They trust documents by inspection: I can see it says that. In a team, inspection matters more. You cannot review, approve or hand over something you cannot see.

More control

Best of both

Captured automatically

Reviewed like documents

Trusted by results and inspection

Documentation

High effort, high control

Trusted by inspection

Less control

Memory

Low effort, low control

Trusted by results

Worst of both

High effort, low control

Hidden and hard to keep

Less effort to keep current

More effort

  • ↑ Memory moves up: show what it knows as readable pages.
  • ← Documentation moves left: agents draft changes, people approve them.
Aim for the empty corner.

The useful corner is the empty one: knowledge captured as easily as memory and checked as easily as a document. Liao's approach gets part of the way, because the agent does the writing. What it lacks is a way to check what the agent wrote. The proposal below aims for the rest.

A proposal: documents backed by evidence​

We have written up a design that keeps Liao's best idea and adds what his approach drops. It is a proposal. None of it is built yet.

The agent reads documents, as Liao suggests. Each statement in a document is a claim, and each claim carries:

  • Citations to the conversations or work that support it.
  • Anchors to what it describes: a file, at a commit, down to the lines.
  • Trust. Whether a person stated it, the agent checked it, or the agent inferred it.
  • Status. Current, suspect, contradicted or replaced.

Claim

Sessions expire after 30 minutes of inactivity.

Page
auth.md › Sessions
Citations
2 conversations, 12 Sep 2026
Anchor
src/auth/session.py @ a1b2c3d, lines 40–58
Trust
Checked by the agent against the code

Suspect: src/auth/session.py changed since a1b2c3d. Queued for re-checking.

An illustrative claim from the proposed design.

The read path changes too:

  • At session start, the agent gets a table of contents. It cannot ask about something it does not know exists, but it can see a map.
  • When the agent opens a file, it gets the claims anchored to that file. The match is exact, not by similarity. It is triggered by what the agent is doing, not what it typed.
  • Before a claim is shown, it is checked against the code. If the file has changed since the claim was anchored, the claim is marked "may be stale" and queued for re-checking.
  • Similarity search stays, as a fallback for everything else.

At the end of a task, the agent proposes changes to the documents, with citations. That is Liao's write-after-work step, with a record of why each change was made.

Documentation loop

operator-memory

  1. Read the relevant documents
  2. Do the work
  3. Update the documents

↺ next task

Anchored documents

Astrocyte design, proposed

  1. Session start: read a table of contents
  2. Open a file: get the claims anchored to it
  3. Check each claim against the code; flag “may be stale”
  4. Do the work
  5. Propose document changes, with citations

↺ next task

The same loop, with a check on whether each claim is still true.

How we would test his claim​

His essay offers no measurement, so the test has to start there.

  1. Build his system as the baseline. Markdown documents only, no retrieval, the agent told to consult and update them.
  2. Test staleness directly. Write memories, change the code underneath them, and measure how often a stale memory misleads the agent. Nobody publishes this today.
  3. Use a shared benchmark. The Agent Memory Leaderboard's coding-memory track compares systems on the same tasks.
  4. Report cost and time with accuracy. Updating documents at the end of every task has a cost too.

Every comparison runs on identical questions, with repeated judging and a margin of error on every number.

What this means for Astrocyte​

Liao's essay changes where Astrocyte is headed. It should sit behind documents as their evidence. Competing with them would be the wrong goal.

  • The prompt hook stops being the default. Recalling by similarity on every prompt is the habit he describes. The plan is to replace it with a table of contents at the start of a session and the claims tied to whatever file the agent opens, keeping similarity as a fallback.
  • Capture stays automatic. Conversations and work are kept as evidence without anyone writing them up. That is the part memory does well.
  • What the agent knows becomes pages people can read. No one should have to query a database to see it.
  • Agents draft, people approve, in the tools the team already uses. Proposed changes go out as pull requests on the repository's Markdown first, then as suggestions in documentation tools. Astrocyte does not need its own review screen.
  • It should work with the documents you already have, including Markdown written by operator-memory, and add sources and staleness checks to them.
  • It should draw on the search systems you already run, treating their results as evidence with sources, never as raw text for the agent to trust.
  • Results come before claims. We will build Liao's system as the baseline and publish the comparison with its margins of error, whichever way it goes.

None of this is built yet. It comes after the benchmark work currently under way, and the design is public for anyone to read and question.

What to ask of any agent memory​

Whether you build on documents, memory or both, three questions separate a system you can trust from one you hope is right:

  • Where did this come from? Every claim the agent reads should point to its source.
  • Is it still true? The system should know when the thing a claim describes has changed.
  • How would you find out it was wrong? If the only way to audit it is a SQL query nobody runs, nobody will.

Liao is right that agents should read documents. They should also be able to check them.