Part 5 of 5 · The loop that learns · How one developer's AI loop remembers, verifies, and enforces
Anyone who uses a coding agent seriously ends up with a setup: a CLAUDE.md or AGENTS.md, some rules, a few skills, a hook or two, a handful of MCP servers. There are now tools to sync that setup across machines, convert it between agents and share it with a team. I wanted to know what people who do this well actually keep in their setups, how they keep them current, and what the evidence says about any of it.
So I looked at about 25 tools and published setups, and six recent papers. Four things stood out.
Vercel ran the cleanest test I found. In their agent evals they gave an agent the same documentation two ways: as an index in AGENTS.md, which is always in context, and as skills the agent could load when it judged them relevant. AGENTS.md passed 100% of cases. Skills passed 53%, the same score as giving the agent no docs at all, and in 56% of cases it never loaded the skill. Telling it explicitly to use skills raised the score to 79%.
Scott Spence measured the same effect from the other side. On its own, Claude Code used a skill for 50 to 55% of his 22 test prompts. With a UserPromptSubmit hook that made it evaluate every skill before starting, activation went to 100%, 22 prompts out of 22.
The lesson for any setup: anything the agent has to decide to look up gets skipped about half the time. If something matters, put it in front of the agent before it starts, or let the harness do the deciding.
That does not mean putting everything in AGENTS.md. Gloaguen and colleagues tested repository context files on SWE-bench tasks and on real issues from repositories that already had them. The files “do not generally improve task success rates”, whether a model or a developer wrote them, while raising inference cost by over 20% on average. The agents did follow the instructions in those files. What did not help were repository overviews, which the authors note are “popular and recommended by model providers”. Their conclusion: context files are useful for specifying non-standard coding practices.
Other results point the same way. IFScale gave models up to 500 instructions at once; the best frontier models followed 68% of them at that density, and favoured the ones that came first. Anthropic’s own advice for CLAUDE.md is blunt: for each line, ask “Would removing this cause Claude to make mistakes?” If not, cut it.
The evidence on file size itself is mixed. Damon McMillan ran 1,650 Claude Code sessions, measuring whether the agent followed one simple convention while he varied file size, instruction position, file architecture and contradictions between files. None of the four had a detectable effect. What did show up was time: each additional function the agent wrote in a session was associated with about 5.6% lower odds of following the convention, though he notes it is not a steady per-step decline. Instructions tend to fade as a session goes on, wherever they sit.
Taken together: keep always-on context to specific, actionable instructions, and expect them to fade during a long session.
The newest evidence is VibeMemBench, which tested memory systems for coding agents on real repository tasks. Injecting past experience whose usefulness had been verified by running the tests raised task resolution by 1.1 to 4.5 percentage points on four of the five agents tested. But when four existing memory systems had to build and retrieve that experience themselves from the same history, 11 of 12 combinations failed to beat the no-memory baseline.
GitHub’s Copilot Memory is the strongest counter-example, and it shows what makes automatic capture work. Its agents save facts about a repository on their own, but each fact carries citations to the code that supports it. Before using a fact, Copilot “checks those citations against the current branch to confirm the information is still accurate. Only validated facts are used.” Facts that go unused for 28 days are deleted. From A/B tests, GitHub reports a 90% merge rate for its coding agent’s pull requests with memory, against 83% without. That is the vendor’s own number.
So the dividing line is not human versus automatic capture. It is whether a memory is checked against the code before the agent relies on it.
ACE, from Zhang and colleagues, adds a point about how memories should change over time. Rewriting a context wholesale leads to “context collapse, where iterative rewriting erodes details over time”; small, structured, incremental updates preserve them. Their approach improved agent benchmarks by 10.6%.
Every setup goes stale, and the most careful ones say so. Aristidis Vasilopoulos built a 108,000-line C# system with coding agents, supported by a 660-line constitution, 19 specialized agents and 34 specification documents, and reports that “specification staleness was the primary failure mode”. When a subsystem changes and its spec does not, “the AI will generate code based on stale information.” His fix is a session-start hook that compares recent commits with a map of subsystems to files, and warns when code changed but its spec did not.
Others do it by hand. One of Sentry’s agent skills ships a SOURCES.md that maps each claim to the file and function supporting it, with the date it was captured and the date part of it was updated. Every’s compound-engineering plugin, the closest thing I found to a learning loop, records solution notes from sessions and has a refresh command that keeps, updates, merges, replaces or deletes them. Its rule is one I would adopt anywhere: “Age alone is not staleness.”
The portability tools solve a different kind of drift. rulesync, Microsoft’s APM and similar tools generate each agent’s configuration from one source, and can fail CI when a generated file no longer matches it (rulesync generate --check exits 1). That catches a stale copy. It does not ask whether the rule itself is still true of the code.
That last point is the gap I have been working on with okl, the subject of this series. Each lesson can be proven by a check you choose, and CI goes red when the code it governs changes, until someone runs the check again. On my own eval, briefing lessons before each task cut repeated known bugs from 43% to 8%, on a small sample of tasks I wrote.
The research also showed me where okl falls short. Agents other than Claude Code have to call for its lessons, which the Vercel result says they often will not. It flags stale lessons in CI rather than at the moment they are used. And it could brief again when a governed file is about to be edited, since instructions fade within a session. Those are next.
I used AI research assistants to survey the tools, published setups and papers, and to check every claim cited here against its source, twice. The checks caught one research summary that misstated a paper’s finding and several smaller overstatements; this piece follows the sources’ own wording. Tools and numbers in this area change weekly; everything here is as of 1 October 2026.
Papers
Measurements and documentation
Setups and tools
--check option