okl — the Observed Knowledge Ledger

A ledger of the lessons a codebase has already paid for, handed to a coding agent before it starts a task, proven by checks and flagged when the code, docs or diagrams they govern change.

PythonSQLite / PostgresFastAPIMCPClaude CodeGitHub ActionsEval harness

Source on GitHub
One task, one model, run twice, in clean copies of the app (receipts in the repo, examples/dotnet-minimal-api). Without the briefing the client sets its own price and the tests still pass; with it, the price comes from the catalogue and the field is gone. The 43% against 8% at the end is the 2026-09-26 run, receipt ab-20260926-0528 in the repo.
Three columns. Left, the plan decided up front: the rules and specs you already wrote, this codebase read by an agent, and outside material audited against your canon, all becoming typed records of symptom, cause and fix. Middle, what reaches the agent before every task: the few records matching this task, action first, failing closed if the store cannot be reached. Right, what a document cannot do: a record names the files it governs, git says when they last changed, and a stale rule fails the pull request that made it stale. Below, the store's type system, and a panel answering why a spec alone does not cover this.

The problem

Two failures, not one, and most knowledge bases answer only the first.

Knowledge that never arrives. A team fixes a subtle bug, learns why it happened, and writes a rule to prevent it. Six weeks later the same class of bug is back, in a different file in the same repository, because the person or agent doing the new work never saw the rule. The knowledge existed. It just wasn’t in front of whoever needed it, at the moment they needed it. Across repositories the same thing costs more: a convention, an architecture decision, a defect class paid for once never reaches the next project, so every repository buys the lesson again at full price.

That failure is older than AI tooling, but agents make it sharper: a model starts every session with no memory of what this codebase has already been burned by, and it writes code faster than anyone can review it back into line.

Knowledge that arrives out of date. The rule is found and applied, but the code it governs moved months ago, so what gets handed over is confident, specific and wrong. The same rot reaches everything a repository says about itself: a README describing a flag that was renamed, an architecture diagram showing a service that was merged away, a figure quoting a percentage from a run whose receipts were never committed. All of it reads as current, because nothing marks it as anything else.

What it does

okl records a lesson once, as a typed, scoped, verified node: a defect class, a piece of canon, a decision, and hands the relevant ones to a coding agent before the first line of code is written. It runs for one repo on day one, and across many when they point at a shared instance, so a lesson learned in one project protects the next one.

The part that’s different

Agent memory is a crowded field: mem0, Zep, Letta, Cognee; Cursor Memories and Devin Knowledge inside the coding agents; AGENTS.md and rules files as the convention standard. Most of them share one property: memories accumulate, and nothing invalidates them.

Several now do invalidate, and those are the honest comparison. GitHub Copilot Memory cites the code lines behind each memory and has the agent re-read them before use. driftlint fails CI when an instruction file names a path, command or import that no longer exists. claude-mem skips a note when the file it came from has changed.

okl goes further for the lessons that matter: it treats a lesson like a test instead of a note. A lesson cites the source it governs, which can be code, a doc or a diagram, and carries a verification receipt from a check you wrote, and okl drift --gate fails CI when that source changes after the lesson was last verified. Re-reading the code, checking that a reference still resolves and decaying on a timer are all judgements made at the moment of use. A receipt is a command that ran and either passed or did not. Knowledge that has quietly gone stale is the failure mode that makes a knowledge base worse than nothing, and it’s the one thing this is built to catch.

Does it work?

Held-fixed A/Bs, with the method and every raw receipt committed in the repo (evals/REPORT.md): 8 authored tasks, 3 samples per arm, generator and blind judge on different models.

What it does not show, stated in the report as plainly as here: the tasks were authored to invite defect classes the store encodes, so it measures what a briefing does when a directly relevant lesson exists. Not general code quality, and not retrieval at scale. n is small. It’s a pilot with receipts, not a benchmark.

Where it came from

Every surface in the encoding loop stops at the repo boundary: canon, tombstones, audits, gates and reviewers each protect the repo they live in and nothing else. This is what carries a lesson past that boundary. The seed corpus is real: dated lessons mined from NextAurora, the riparian pipeline, and the Quartzose eval work, including the ones that cost the most to learn.

Running on this site

This site is built with okl wired in, so it is a user of the tool and not only a page about it. A prompt-submit hook briefs every coding session before work starts, and it fails closed: if the store cannot be reached, the prompt is blocked rather than answered with no lessons. The store holds 16 lessons from building this site, 12 of them org-scoped because nothing about them is particular to it: a CSP that silently blocked the webfonts in production while the dev server showed them fine, an Astro whitespace rule that deletes spaces next to tags, a grid minimum that overflows a phone.

Running it here also found three defects in okl itself. The install instructions said pip install okl, which cannot work because PyPI refuses that name, and the tag vocabulary was compiled into the package, so this repo’s frontend and prose lessons would not import; both shipped in 0.5.0. The third was worse: the hooks resolved their config by walking up from the working directory, so a session whose shell had moved to a scratch folder had every prompt blocked as “unreachable” against a store that was fine, and the message named connectivity rather than the cause. They anchor to the project directory from 0.6.0.

Status

pipx install observed-knowledge-ledger installs it, or install it into Claude Code as a plugin (/plugin marketplace add emeraldleaf/okl); either way the command is okl. The distribution name and the command differ because PyPI refuses okl as confusable.

It is at 0.7.0. Since the measurements above it has grown the things a tool needs before anyone else can run it: okl doctor, which names a repo wired twice or colliding with another memory tool; an installer that fingerprints every file it writes, so an upgrade never overwrites an edit, and an --uninstall that removes its own files and nothing else; and two worked examples, a FastAPI service and a .NET one, each playing the bait tasks against a control copy and a briefed copy with the diffs kept.

Still a v0: flat retrieval by decision (an ADR sets the threshold at which that stops being enough), a small store, one author. It is published because the bet is legible and the receipts are checkable, not because it is finished.