The encoding loop

Agentic coding tools ship with opinions baked in from training. Left unconstrained they apply those defaults, quietly, confidently, every session. The loop below turns each catch into a rule the agent loads next time, on the lightest surface that holds it.

The problem it solves

A code review catches a bad pattern. You fix it. Six sessions later the same pattern comes back, because nothing about that conversation persisted: not for the agent, not for the next person, not for future-you. The fix lived in a PR; the rule lived nowhere.

The encoding loop makes that a defect. A merged fix without its corresponding encoding is a half-finished job: the same session that surfaces a non-obvious failure also writes it down, on exactly one surface, the smallest one that holds it. The threshold is a single question: if the next person could repeat this mistake, or would have to re-derive the rule from first principles, it belongs in writing.

This is deliberately the dual of spec-driven development. Spec-driven encodes per-feature requirements up front and regenerates code from them. This encodes codebase-wide invariants once, and every feature after it starts smarter. Roughly zero upfront cost per feature; the encoding cost amortizes across every future session. Its failure mode is honest too: rule sprawl, and encoding fatigue.

Five surfaces, ordered by enforcement strength

Rules are encoded on the softest surface that works, and promoted down the spectrum only when they earn it: when the soft version keeps getting missed.

#SurfaceWhat goes hereCatches violations at…
1Always-on canon
CLAUDE.md
Invariants only, one line each, with a link to the deeper doc. Loaded into every session, so every byte is cognitive overhead for the model and the human.Code-generation time
2Deep reference
docs/ + paired diagrams
The why behind a rule, and the picture reviewers reason against. Read on demand, not every session.Review and onboarding
3Procedure rituals
.claude/ commands, agents, skills
Multi-step motions worth a single keystroke, and a reviewer that starts with fresh context instead of the conversation that produced the code.On invocation, before code lands
4PR automation
.coderabbit.yaml path_instructions
File-glob-scoped guidance that re-checks a rule at review time without re-deriving it. Non-blocking: a reviewer can dismiss it.Every diff, automatically
5Mechanical gate
CI scripts, analyzers, tests, hooks
The build fails. Nobody has to notice, remember, or agree.Merge time, loudly

Those five cluster into three tiers, and the tiers are the point. Convention (1–2) relies on someone noticing, and noticing is the first thing that degrades under deadline. PR automation (3–4) fires on its own but can be dismissed. Mechanical gates (5) fail loudly and need no human in the loop at all. Most rules should never reach tier three. The ones that do are the ones that already bit twice.

All five stop at the repo boundary. A lesson encoded here protects this repo and nothing else, and that limit is what okl exists to answer: a ledger the surfaces are mined into, read before a task starts and written back to after, fail-closed, so an unreachable store is never mistaken for "no lessons apply". It runs in this site, in okl itself and in Quartzose. Scope is the whole trick: a lesson marked for the org is meant for every repo, one marked for a repo never leaves it. Each of the three keeps its own store today, so a shared lesson travels when its seed file is loaded, and the shared service okl ships is not yet pointed at. In a held-fixed A/B with committed receipts (48 runs per comparison, blind cross-model judging), defect reproduction with the briefing has stayed at 4–13% while the unbriefed baseline read 33–50% across runs; in the first runs a briefed budget model (12%) also beat the unbriefed frontier model (33%), a result not yet repeated. Method, limits and raw receipts are in the repo (evals/REPORT.md), and the story is told in the five-part series.

Wired into a real repo

Everything below is checked into NextAurora, the five-service .NET event-driven system you can break live. It is the working instance of the method, not a description of one.

Edit time: .claude/

A second AI reviewer (GitHub Copilot, in-editor) runs against its own instruction file. The operating principle: disagreement between two models is a signal to dig, not a vote to settle.

Build time: the analyzer wall

TreatWarningsAsErrors and AnalysisMode=All. Zero warnings allowed, so every rule below is a build break, not a suggestion.

Test time: Testcontainers, not mocks

Integration tests run against real engines in Docker. That choice is what makes the rest of the architecture affordable: with a real database one command away, the repository wrapper that exists only to be mocked stops earning its keep, and handlers can take the DbContext directly.

SliceContainersWhat it proves
CatalogPostgres + RedisCache invalidation, xmin concurrency, gRPC server, search projection
OrderSQL Server + RabbitMQOutbox staging, rowversion concurrency, saga publish side, read projection, durable inbox over a real broker
PaymentSQL ServerAcceptor→gateway split, outbox atomicity, idempotency
ShippingPostgresIDOR-safe read predicate, saga consume side, idempotency under at-least-once delivery

Three test patterns are required, not encouraged, and every review pass checks for them:

Architecture tests (NetArchTest) hold the line that a four-project Clean Architecture split would otherwise hold: every service’s Domain/ namespace is asserted to have no dependency on EF Core, ASP.NET, the message bus, or any data provider. Same boundary, one test-suite slot, none of the ceremony. The full argument is written up here.

PR time: three reviewers, three lenses

CI itself runs three parallel jobs: build plus unit tests, the frontend (lint, typecheck, Vitest with Testing Library and MSW, build), and the Testcontainers integration suite in its own job so a fast unit signal is never gated on container startup. Both .NET jobs upload Cobertura coverage to Codecov under separate unit and integration flags, and test results surface as an annotated PR check rather than a buried job log. Coverage gates only on regression: total coverage falling by more than a point fails the check, while coverage of the lines a PR adds is posted for review and never blocks. The number is a floor against backsliding, not a claim about quality.

The guards nobody ships off the shelf

These are the interesting ones. Each exists because something specific went wrong once, and each is a short script run in CI: the smallest mechanism that closes the hole.

Tombstone audit

Identifiers removed from the code (a retired transport’s API names, dead queue names, old metric names) must not survive in prose as if they were current. The compiler catches stale identifiers in code; nothing catches them in docs.

A messaging-transport swap left 15+ docs teaching the removed transport as current. Found by a human reading, not by a machine.

Source →

Broken-link audit

Every relative markdown link to a local source file must resolve. Skips external URLs and anchors.

An architecture refactor deleted files that CLAUDE.md and the demo docs still cited by path.

Source →

Doc-path audit

Every backtick-quoted repo path and every #L line anchor in the markdown must resolve. The broken-link audit only sees markdown links; this covers the paths written inline.

A drift sweep found 15 inline paths and 8 line anchors that had rotted past the link audit. Each one was a doc pointing a reader, or an agent, at code that had moved.

Source →

Broken-COPY audit

Every COPY/ADD source path in every Dockerfile must exist in the build context.

A Dockerfile rotted silently for months after a refactor. The broken paths would only have failed at deploy time, so a live demo kept serving pre-refactor code.

Source →

Diagram-pair audit

Every diagram exists as both an editable .excalidraw source and a rendered .svg. One without the other is a broken diagram: an SVG nobody can edit, or a source nobody can see on github.com.

Diagrams are the review surface, not a byproduct. Treating them as build artifacts is what keeps them true.

Source →

CLAUDE.md size budget

Warns at 400 lines, fails at 500. Over budget, detail moves to docs/ or a skill and the canon keeps a headline plus a link.

The failure mode of every rule-encoding method is decay into the spec it replaced. A hard line count is the cheapest defense that actually holds.

Source →

Static mutable collection audit

Bans plain List/Dictionary/HashSet/Queue/Stack as static fields. ConcurrentDictionary, ImmutableList, FrozenDictionary and Channel pass.

The one concurrency hazard an analyzer can’t catch: the difference between an immutable lookup table and shared mutable state is structural, not syntactic.

Source →

Messaging-topology audit

Every queue constant reaches a publisher-side declaration; every listener is durable-inbox or inline; no inline topology literals.

Most integration tests stub the transport, so they cannot catch a queue that is declared but never bound.

Source →

Considered and skipped

A tooling list that only grows is a smell. These were evaluated and left out, on the record:

ToolWhy not
SonarQube (self-hosted) and SonarCloud (hosted)The detection already runs at build time via SonarAnalyzer.CSharp, the same engine. A dashboard would add hosting overhead and a second place to look without adding findings.
Spec Kit: /specify + /plan + /tasksPer-feature spec authoring converts a compounding advantage into ceremony. Same lineage, opposite end of the stick.
MCP servers for project contextThe canon, the paired docs and the greppable surfaces already give the agent persistent decision context, and they live in the repo, next to the code they describe. An external tool layer over decisions that already live in-repo is convergence in the wrong direction. OKL does ship an MCP server, and the line it draws is the same one: it carries only what has to cross a repo boundary, which no in-repo surface can.
A CI-generator skillThe CI works. Generating it again is anti-pragmatic.

The decay problem

Every method eventually decays toward the thing it replaced. Agile decayed into spec-process-with-shorter-cycles. Spec-driven development started as goal-direction and decayed into 1,300-line specs. TDD decayed into coverage-gaming. There is no reason to assume this one is exempt, so the defenses are built in: a hard size budget on the canon so it can’t bloat into the spec it replaced; a drift audit so paraphrases stay aligned with the source; an explicit separation between lessons from this feature and what this feature must do, so the rule set never collapses into per-feature ceremony; and a promotion spectrum, so the rules that matter become loud failures rather than quiet conventions.

The claim isn’t that AI made the work faster. A 2025 controlled trial found experienced developers were measurably slower with AI while feeling faster. Being wrong while feeling fast is the whole failure in one sentence. The claim is narrower and testable: the build eventually enforces what used to be a fifteen-minute PR review conversation, and that conversation never has to happen again.

Read further