Agentic coding tools ship with opinions baked in from training. Left unconstrained they apply those defaults, quietly, confidently, every session. The loop below turns each catch into a rule the agent loads next time, on the lightest surface that holds it.
A code review catches a bad pattern. You fix it. Six sessions later the same pattern comes back, because nothing about that conversation persisted: not for the agent, not for the next person, not for future-you. The fix lived in a PR; the rule lived nowhere.
The encoding loop makes that a defect. A merged fix without its corresponding encoding is a half-finished job: the same session that surfaces a non-obvious failure also writes it down, on exactly one surface, the smallest one that holds it. The threshold is a single question: if the next person could repeat this mistake, or would have to re-derive the rule from first principles, it belongs in writing.
This is deliberately the dual of spec-driven development. Spec-driven encodes per-feature requirements up front and regenerates code from them. This encodes codebase-wide invariants once, and every feature after it starts smarter. Roughly zero upfront cost per feature; the encoding cost amortizes across every future session. Its failure mode is honest too: rule sprawl, and encoding fatigue.
Rules are encoded on the softest surface that works, and promoted down the spectrum only when they earn it: when the soft version keeps getting missed.
| # | Surface | What goes here | Catches violations at… |
|---|---|---|---|
| 1 | Always-on canonCLAUDE.md | Invariants only, one line each, with a link to the deeper doc. Loaded into every session, so every byte is cognitive overhead for the model and the human. | Code-generation time |
| 2 | Deep referencedocs/ + paired diagrams | The why behind a rule, and the picture reviewers reason against. Read on demand, not every session. | Review and onboarding |
| 3 | Procedure rituals.claude/ commands, agents, skills | Multi-step motions worth a single keystroke, and a reviewer that starts with fresh context instead of the conversation that produced the code. | On invocation, before code lands |
| 4 | PR automation.coderabbit.yaml path_instructions | File-glob-scoped guidance that re-checks a rule at review time without re-deriving it. Non-blocking: a reviewer can dismiss it. | Every diff, automatically |
| 5 | Mechanical gateCI scripts, analyzers, tests, hooks | The build fails. Nobody has to notice, remember, or agree. | Merge time, loudly |
Those five cluster into three tiers, and the tiers are the point. Convention (1–2) relies on someone noticing, and noticing is the first thing that degrades under deadline. PR automation (3–4) fires on its own but can be dismissed. Mechanical gates (5) fail loudly and need no human in the loop at all. Most rules should never reach tier three. The ones that do are the ones that already bit twice.
All five stop at the repo boundary. A lesson encoded here protects this repo and nothing else, and that limit is what okl exists to answer: a ledger the surfaces are mined into, read before a task starts and written back to after, fail-closed, so an unreachable store is never mistaken for "no lessons apply". It runs in this site, in okl itself and in Quartzose. Scope is the whole trick: a lesson marked for the org is meant for every repo, one marked for a repo never leaves it. Each of the three keeps its own store today, so a shared lesson travels when its seed file is loaded, and the shared service okl ships is not yet pointed at. In a held-fixed A/B with committed receipts (48 runs per comparison, blind cross-model judging), defect reproduction with the briefing has stayed at 4–13% while the unbriefed baseline read 33–50% across runs; in the first runs a briefed budget model (12%) also beat the unbriefed frontier model (33%), a result not yet repeated. Method, limits and raw receipts are in the repo (evals/REPORT.md), and the story is told in the five-part series.
Everything below is checked into NextAurora, the five-service .NET event-driven system you can break live. It is the working instance of the method, not a description of one.
.claude/PreToolUse hook rejects .Result / .Wait() / .GetAwaiter().GetResult() in proposed edits. The build-time analyzer catches the same patterns later. The hook catches them earlier, so the bad diff never exists. Another fires on every edit to the canon and lists the files that paraphrase it, so drift gets reviewed the same session./check-rules, which audits every paraphrase of a canon rule against the canonical text and flags drift, and /drift-audit, which checks every doc, comment and canon rule against the code as it is now. Its first run confirmed 137 findings after every mechanical gate had passed.A second AI reviewer (GitHub Copilot, in-editor) runs against its own instruction file. The operating principle: disagreement between two models is a signal to dig, not a vote to settle.
TreatWarningsAsErrors and AnalysisMode=All. Zero warnings allowed, so every rule below is a build break, not a suggestion.
BannedSymbols.txt: Task.WaitAll, Parallel.For, Thread.Sleep and friends, each with custom replacement guidance in the error message.Integration tests run against real engines in Docker. That choice is what makes the rest of the architecture affordable: with a real database one command away, the repository wrapper that exists only to be mocked stops earning its keep, and handlers can take the DbContext directly.
| Slice | Containers | What it proves |
|---|---|---|
| Catalog | Postgres + Redis | Cache invalidation, xmin concurrency, gRPC server, search projection |
| Order | SQL Server + RabbitMQ | Outbox staging, rowversion concurrency, saga publish side, read projection, durable inbox over a real broker |
| Payment | SQL Server | Acceptor→gateway split, outbox atomicity, idempotency |
| Shipping | Postgres | IDOR-safe read predicate, saga consume side, idempotency under at-least-once delivery |
Three test patterns are required, not encouraged, and every review pass checks for them:
GET /orders/{id} IDOR survived the lifetime of the codebase.Architecture tests (NetArchTest) hold the line that a four-project Clean Architecture split would otherwise hold: every service’s Domain/ namespace is asserted to have no dependency on EF Core, ASP.NET, the message bus, or any data provider. Same boundary, one test-suite slot, none of the ceremony. The full argument is written up here.
request_changes_workflow on and 23 path-scoped instruction globs: VSA slice shape, aggregate factory rules, endpoint authorization and the 404-not-403 shape, the outbox-outside-a-handler trap, migration immutability, test conventions, middleware order. That file is surface #4: it is how a rule written once in the canon gets re-checked at review time without anyone re-deriving it.security-and-quality query set), plus Gitleaks secret scanning that blocks the merge, and Dependabot on a weekly NuGet cadence.CI itself runs three parallel jobs: build plus unit tests, the frontend (lint, typecheck, Vitest with Testing Library and MSW, build), and the Testcontainers integration suite in its own job so a fast unit signal is never gated on container startup. Both .NET jobs upload Cobertura coverage to Codecov under separate unit and integration flags, and test results surface as an annotated PR check rather than a buried job log. Coverage gates only on regression: total coverage falling by more than a point fails the check, while coverage of the lines a PR adds is posted for review and never blocks. The number is a floor against backsliding, not a claim about quality.
These are the interesting ones. Each exists because something specific went wrong once, and each is a short script run in CI: the smallest mechanism that closes the hole.
Identifiers removed from the code (a retired transport’s API names, dead queue names, old metric names) must not survive in prose as if they were current. The compiler catches stale identifiers in code; nothing catches them in docs.
A messaging-transport swap left 15+ docs teaching the removed transport as current. Found by a human reading, not by a machine.
Source →Every relative markdown link to a local source file must resolve. Skips external URLs and anchors.
An architecture refactor deleted files that CLAUDE.md and the demo docs still cited by path.
Source →Every backtick-quoted repo path and every #L line anchor in the markdown must resolve. The broken-link audit only sees markdown links; this covers the paths written inline.
A drift sweep found 15 inline paths and 8 line anchors that had rotted past the link audit. Each one was a doc pointing a reader, or an agent, at code that had moved.
Source →Every COPY/ADD source path in every Dockerfile must exist in the build context.
A Dockerfile rotted silently for months after a refactor. The broken paths would only have failed at deploy time, so a live demo kept serving pre-refactor code.
Source →Every diagram exists as both an editable .excalidraw source and a rendered .svg. One without the other is a broken diagram: an SVG nobody can edit, or a source nobody can see on github.com.
Diagrams are the review surface, not a byproduct. Treating them as build artifacts is what keeps them true.
Source →Warns at 400 lines, fails at 500. Over budget, detail moves to docs/ or a skill and the canon keeps a headline plus a link.
The failure mode of every rule-encoding method is decay into the spec it replaced. A hard line count is the cheapest defense that actually holds.
Source →Bans plain List/Dictionary/HashSet/Queue/Stack as static fields. ConcurrentDictionary, ImmutableList, FrozenDictionary and Channel pass.
The one concurrency hazard an analyzer can’t catch: the difference between an immutable lookup table and shared mutable state is structural, not syntactic.
Source →Every queue constant reaches a publisher-side declaration; every listener is durable-inbox or inline; no inline topology literals.
Most integration tests stub the transport, so they cannot catch a queue that is declared but never bound.
Source →A tooling list that only grows is a smell. These were evaluated and left out, on the record:
| Tool | Why not |
|---|---|
| SonarQube (self-hosted) and SonarCloud (hosted) | The detection already runs at build time via SonarAnalyzer.CSharp, the same engine. A dashboard would add hosting overhead and a second place to look without adding findings. |
Spec Kit: /specify + /plan + /tasks | Per-feature spec authoring converts a compounding advantage into ceremony. Same lineage, opposite end of the stick. |
| MCP servers for project context | The canon, the paired docs and the greppable surfaces already give the agent persistent decision context, and they live in the repo, next to the code they describe. An external tool layer over decisions that already live in-repo is convergence in the wrong direction. OKL does ship an MCP server, and the line it draws is the same one: it carries only what has to cross a repo boundary, which no in-repo surface can. |
| A CI-generator skill | The CI works. Generating it again is anti-pragmatic. |
Every method eventually decays toward the thing it replaced. Agile decayed into spec-process-with-shorter-cycles. Spec-driven development started as goal-direction and decayed into 1,300-line specs. TDD decayed into coverage-gaming. There is no reason to assume this one is exempt, so the defenses are built in: a hard size budget on the canon so it can’t bloat into the spec it replaced; a drift audit so paraphrases stay aligned with the source; an explicit separation between lessons from this feature and what this feature must do, so the rule set never collapses into per-feature ceremony; and a promotion spectrum, so the rules that matter become loud failures rather than quiet conventions.
The claim isn’t that AI made the work faster. A 2025 controlled trial found experienced developers were measurably slower with AI while feeling faster. Being wrong while feeling fast is the whole failure in one sentence. The claim is narrower and testable: the build eventually enforces what used to be a fifteen-minute PR review conversation, and that conversation never has to happen again.