A ledger of the lessons a codebase has already paid for, handed to a coding agent before it starts a task, proven by checks and flagged when the code, docs or diagrams they govern change.
PythonSQLite / PostgresFastAPIMCPClaude CodeGitHub ActionsEval harness
Two failures, not one, and most knowledge bases answer only the first.
Knowledge that never arrives. A team fixes a subtle bug, learns why it happened, and writes a rule to prevent it. Six weeks later the same class of bug is back, in a different file in the same repository, because the person or agent doing the new work never saw the rule. The knowledge existed. It just wasn’t in front of whoever needed it, at the moment they needed it. Across repositories the same thing costs more: a convention, an architecture decision, a defect class paid for once never reaches the next project, so every repository buys the lesson again at full price.
That failure is older than AI tooling, but agents make it sharper: a model starts every session with no memory of what this codebase has already been burned by, and it writes code faster than anyone can review it back into line.
Knowledge that arrives out of date. The rule is found and applied, but the code it governs moved months ago, so what gets handed over is confident, specific and wrong. The same rot reaches everything a repository says about itself: a README describing a flag that was renamed, an architecture diagram showing a service that was merged away, a figure quoting a percentage from a run whose receipts were never committed. All of it reads as current, because nothing marks it as anything else.
okl records a lesson once, as a typed, scoped, verified node: a defect class, a piece of
canon, a decision, and hands the relevant ones to a coding agent before the first
line of code is written. It runs for one repo on day one, and across many when they point
at a shared instance, so a lesson learned in one project protects the next one.
Agent memory is a crowded field: mem0, Zep, Letta, Cognee; Cursor Memories and Devin
Knowledge inside the coding agents; AGENTS.md and rules files as the convention standard.
Most of them share one property: memories accumulate, and nothing invalidates them.
Several now do invalidate, and those are the honest comparison. GitHub Copilot Memory cites
the code lines behind each memory and has the agent re-read them before use. driftlint
fails CI when an instruction file names a path, command or import that no longer exists.
claude-mem skips a note when the file it came from has changed.
okl goes further for the lessons that matter: it treats a lesson like a test instead of a
note. A lesson cites the source it governs, which can be code, a doc or a diagram, and
carries a verification receipt from a check you wrote, and okl drift --gate fails CI when
that source changes after the lesson was last verified. Re-reading the code, checking that a
reference still resolves and decaying on a timer are all judgements made at the moment of
use. A receipt is a command that ran and either passed or did not. Knowledge that has
quietly gone stale is the failure mode that makes a knowledge base worse than nothing, and
it’s the one thing this is built to catch.
Held-fixed A/Bs, with the method and every raw receipt committed in the repo
(evals/REPORT.md): 8 authored tasks, 3 samples per arm, generator and blind judge on
different models.
n, so
movement inside it is not a result. The briefing gap, 33–50% against 4–13%, clears it.
Every run has receipts.What it does not show, stated in the report as plainly as here: the tasks were authored to
invite defect classes the store encodes, so it measures what a briefing does when a directly
relevant lesson exists. Not general code quality, and not retrieval at scale. n is small.
It’s a pilot with receipts, not a benchmark.
Every surface in the encoding loop stops at the repo boundary: canon, tombstones, audits, gates and reviewers each protect the repo they live in and nothing else. This is what carries a lesson past that boundary. The seed corpus is real: dated lessons mined from NextAurora, the riparian pipeline, and the Quartzose eval work, including the ones that cost the most to learn.
This site is built with okl wired in, so it is a user of the tool and not only a page
about it. A prompt-submit hook briefs every coding session before work starts, and it fails
closed: if the store cannot be reached, the prompt is blocked rather than answered with no
lessons. The store holds 16 lessons from building this site, 12 of them org-scoped because
nothing about them is particular to it: a CSP that silently blocked the webfonts in
production while the dev server showed them fine, an Astro whitespace rule that deletes
spaces next to tags, a grid minimum that overflows a phone.
Running it here also found three defects in okl itself. The install instructions said
pip install okl, which cannot work because PyPI refuses that name, and the tag vocabulary
was compiled into the package, so this repo’s frontend and prose lessons would not
import; both shipped in 0.5.0. The third was worse: the hooks resolved their config by
walking up from the working directory, so a session whose shell had moved to a scratch
folder had every prompt blocked as “unreachable” against a store that was fine, and the
message named connectivity rather than the cause. They anchor to the project directory
from 0.6.0.
pipx install observed-knowledge-ledger installs it, or install it into Claude Code as a
plugin (/plugin marketplace add emeraldleaf/okl); either way the command is okl. The
distribution name and the command differ because PyPI refuses okl as confusable.
It is at 0.7.0. Since the measurements above it has grown the things a tool needs before
anyone else can run it: okl doctor, which names a repo wired twice or colliding with
another memory tool; an installer that fingerprints every file it writes, so an upgrade
never overwrites an edit, and an --uninstall that removes its own files and nothing
else; and two worked examples, a FastAPI service and a .NET one, each playing the bait
tasks against a control copy and a briefed copy with the diffs kept.
Still a v0: flat retrieval by decision (an ADR sets the threshold at which that stops being enough), a small store, one author. It is published because the bet is legible and the receipts are checkable, not because it is finished.