The first time we raced our Belgian encoding against EUROMOD — ran the two implementations of the same law head-to-head over the same simulated population, comparing every number they produced — the two disagreed by €643. The encoding had computed roughly €20.9 billion in employee and employer social-security contributions across a simulated population of workers; EUROMOD — the European Union's reference tax-benefit model, decades of institutional work deep [1] — came out €643 higher. Six hundred forty-three euros, on twenty-one billion. Is that success?
The honest answer is that the number alone cannot say, because €643 is consistent with several different worlds. It could be floating-point dust — hundreds of thousands of sub-cent rounding differences, summed. It could be a single catastrophic error: one worker off by €643 while every other matches to the euro. It could be a small systematic bias, a rule applied a hair too generously to everyone. Or it could be the worst case, the one that looks best from a distance: large errors in opposite directions, canceling almost perfectly on their way to a reassuring total. This book has already met that case. When CBO scored the Affordable Care Act in 2010, its projection of the total uninsured rate landed near the outcome while — as later decomposition showed — overstating exchange enrollment by more than half and understating Medicaid by nearly as much, two errors partly canceling. The total was approximately right for the wrong reasons. If the only test is whether two totals sit close together, you cannot distinguish a right model from a flatteringly wrong one.
So we did not test the totals. We compared workers. For each of the 78,479 simulated workers, the pipeline lined up the encoding's contribution against EUROMOD's and took the difference, and the largest single-worker discrepancy in the entire population was three cents a year. Not on average — three cents at the very worst. Every worker matched to within a rounding error, and the €643 was exactly what it looked like once examined unit by unit: dust.
Aggregate agreement can hide anything. Per-unit comparison hides nothing.
That sentence is the book's discipline compressed, and this chapter is its mechanics. Chapter 9 ended by conceding that an encoding can compile, pass its tests, and ground every dollar in quoted statute while still being wrong about the world. The question of whether to trust law encoded by machines comes down to whether that last gap can be closed — and the answer is that "trust" is the wrong frame. You do not trust an encoding. You verify it, one unit at a time, against something that is true independently of the machine, and you wire the verification in so tightly that a failing encoding cannot ship. An encoding no one has checked, determination by determination, against ground truth is not a cheaper version of a validated model. It is confident fiction with good production values.
The ladder of checks
The ladder starts where chapter 9 left off, with the money-atom gate — merge-blocking, remember, not advisory. Every rate, threshold, taper, and cap that actually moves money must be grounded in a quoted excerpt from the governing source; an amount with no provision behind it fails the build the way a syntax error does. The finished Belgian encoding carried 539 distinct monetary obligations, and the gate demanded a source excerpt for every one; the count of ungrounded obligations stood at zero. The interesting part is not that the count reached zero but that the gate is configured so it can only ever fall.
Grounding is necessary and nowhere near sufficient. A number can be faithfully quoted from the statute and wired into the wrong formula, or into the right formula with the wrong sign. So the encoding must also compile, pass the companion tests that ship beside every rule, and clear a proof pass that checks the logic holds together — that its pieces refer to quantities that exist, in units that match, along paths that resolve. All of that establishes internal coherence and external sourcing. None of it establishes correctness, because a system can be coherent, sourced, and wrong.
For correctness, the encoding is raced against an oracle: an independent implementation of the same law, usually built by other people, often over many years, whose answers function as a second opinion we did not write and cannot quietly influence. The oracles in use are the ones the field already trusts. TAXSIM, the NBER calculator that has anchored academic tax research for four decades .[2] The classic PolicyEngine engine. UKMOD for the United Kingdom .[3] EUROMOD for EU member states .[1] The SOUTHMOD family that UNU-WIDER maintains for a growing list of developing economies .[4] And wherever a tax authority publishes an official calculator, the authority's own.
Chapter 3 laid out three levels on which any model can be validated — component, aggregate, predictive. The oracle program industrializes the first. Component validation used to mean spot-checking one case against an IRS worksheet; here it means replaying entire populations through two independent engines and demanding agreement on every record. Component validation grown teeth. And the whole-population part is load-bearing: a hand-written test suite checks only the situations its author imagined — the married couple, the pensioner, the three-child family — and says nothing about the case nobody imagined, the worker crossing two thresholds in the same year, the household in the seam between two rules. A full calibrated population contains those cases whether or not anyone anticipated them. Racing it through both engines tests them too.
What parity looks like
A word about whose numbers the rest of this section reports: ours. The conformance boards are public and the comparisons are built to be rerun, but what follows is our program measuring its own encodings against reference models other people built — a builder's report, not an outside audit. That is precisely why the machinery has to be this strict.
The United Kingdom is where the encodings and the oracle agree most exactly, and "exactly" is the operative word. For a deterministic calculation there is a single right answer, and a tolerance band is mostly a comfortable place for a wrong answer to hide. UK income tax matches UKMOD to the pound — not within a tolerance — across five test suites, including the withdrawal of the personal allowance, the taper that claws back the tax-free band from high earners and produces the marginal-rate spike that trips up informal tax advice. The Scottish income tax, with its own six-band structure, matches on nine cases out of nine. Savings income and dividend income, each taxed on its own schedule with its own allowances, match exactly. National Insurance matches or is explained: where the two systems diverge, the divergence traces to a documented convention — UKMOD computes contributions weekly and then annualizes, while the encoding works in annual terms — and once you account for it, the two reconcile.
Universal Credit is the instructive case, because the raw score looks bad: of nineteen test cases, eight matched. The other eleven are not errors. UKMOD applies a stochastic take-up draw — deciding at random, household by household, whether an eligible family actually claims, because in reality many eligible families never do — while the encoding computes statutory entitlement, what the household is owed if it claims. Those are different quantities by construction, and every one of the eleven mismatches resolves precisely to that difference. The two systems are answering two different questions, and both are answering theirs correctly.
Which is why the load-bearing word in this program is not "match" but disposition. A raw mismatch is neither a failure nor a pass; it is an open question, and it stays open until someone closes it with evidence. An explanation, here, is a checked artifact rather than a shrug: an arithmetic reconciliation that reproduces the gap to the penny, a citation to the statutory provision the oracle implements differently, or a documented account of the oracle's own modeling convention, like the take-up draw. Every mismatch is either dispositioned or unexplained, and only unexplained mismatches count against the encoding. The standing UK result — the number that matters, not the raw match rate — is that across every surface we have compared, 100 percent of mismatches are explained and zero are unexplained.
Belgium tells the same story with different furniture. The regional and municipal surcharges stacked on Belgian income tax match nine cases out of nine. A greenfield encoding of student study allowances — a body of rules with no prior implementation anywhere to copy from — matches six out of six. The contributions match per worker to three cents. Every mismatch on the Belgian conformance board carries its disposition, and none is attributed to an error on the encoding's side. "Clean," on these boards, means everything that disagreed was run to ground and written down.
When the oracle is wrong
The dispositions discipline has a consequence we did not design for. If every mismatch must be chased to evidence, then sometimes the evidence exonerates the encoding and implicates the reference model. Race a fresh, exhaustively grounded implementation against a model maintained by hand for decades, and now and then the fresh one is right where the mature one is wrong.
Exhaustiveness makes that possible. Nobody maintaining a mature model by hand replays tens of thousands of records against a second independent engine on every change; until recently the compute and the tooling made it prohibitive. When you do, an inconsistency that surfaces in a handful of unusual cases stops hiding in the tail. It becomes a specific record with a specific wrong number that a person can investigate.
Chasing one mismatch through the UK savings rules, we found the discrepancy was not in our encoding. UKMOD was mishandling an interaction involving the personal savings allowance, a corner where two provisions meet and the combined behavior is easy to get subtly wrong. We wrote it up the way the discipline demands — the inputs, the statutory result expected and why, the reference model's actual output, the provision each system implemented — and filed it publicly as an issue on the UKMOD-PUBLIC repository for its maintainers to weigh .[5] A second case was stranger. Running the EUROMOD comparison, identical input cases sometimes scored differently depending on where they sat in a batch — the same household, the same year, two answers, determined by the case's neighbors in the queue rather than by anything about the case. That is the fingerprint of state leaking between records that are supposed to be independent. We root-caused it to a batch-processing contamination bug in the adapter layer that runs cases through the model, and reported it to the maintainers with a reproduction in hand.
These are not gotcha stories, and the framing matters more than the finding. The reference models are why our encodings can be trusted at all; strip away the independent second opinion and "the machine says so" is exactly the epistemic position this project exists to escape. TAXSIM has run for forty years. EUROMOD and its national variants represent decades of institutional work. That a model of that size contains a bug is not a scandal — every large model does, our encodings emphatically included — and finding one is no victory over its authors. The collaboration runs in both directions: the reference model gets free, specific, evidence-backed quality assurance from a system precise enough to notice a three-cent gap and disciplined enough to chase it to cause. Openness means the errors flow both ways, bugs filed upstream and downstream, both models better for it.
A one-way valve
The hardest honest objection to a public rules layer is rot. Laws change every year, and the characteristic failure of shared infrastructure is quiet staleness — encodings drifting out of date, coverage eroding at the edges, a model accurate on the day it shipped and subtly wrong three budgets later with no one the wiser. A verification suite that runs once and gets admired is no defense. A suite wired into a ratchet is.
The conformance gates are configured as one-way valves. Coverage — the number of statutory surfaces under test against an oracle — may only rise; a change that would drop a jurisdiction or program below its current tested coverage fails CI and cannot merge. The count of unexplained mismatches may only fall; a change that introduces a new unexplained divergence, or quietly un-explains a previously dispositioned one, fails too. The count of ungrounded monetary obligations may only fall toward the zero it already reached. The valves hold because the full conformance suite runs on every change, not on a schedule someone has to remember. When a new budget moves a threshold, the new value lands as a dated addition beside the old one and must pass against the oracle for its own year before merging — last year's cases checked against last year's law, this year's against this year's. Nothing is overwritten in place, so nothing quietly falls out of coverage. The moment the encodings drift from the law is the moment the build stops going green.
Two honest limits. A rule can still go un-encoded, and a brand-new program can still escape coverage entirely; the ratchet preserves what has been won and checked, and it does not discover law it has never seen. What it converts is the failure mode: staleness stops being an anxious, permanently-behind human maintenance problem and becomes a mechanical invariant on the tested surface — correctness, once won and checked, cannot be lost again without the build going red, and silence was always the dangerous part. That promise is weaker than "always current." It is much stronger than "trust us to keep it current."
The gates are for us first
It would be comfortable to describe all of this as a wall against untrustworthy AI. That is true, and it is not the honest emphasis. The gates catch our own mistakes first, and the most useful report I can give is the times they turned red on us.
The B1 probe from chapter 9 is the cleanest example, and it is worth restating as a verification story rather than an engineering one. We took twenty-five modules in production, regenerated each fresh from its source, and compared the output against the version we were running — regeneration cost about eight cents of compute per module, and twenty-three of the twenty-five came back with the old version correct and the regenerated one worse, mostly through renamed variables that would have silently broken every downstream consumer. Nobody reasoned their way to that in advance; we know it because a gate compared regenerated output to a known-good baseline, and the baseline won, 23 to 2. The deeper lesson outlived the experiment: an encoding is an answer with a history — which source, which version of the law, which call was made where the text was genuinely ambiguous — and regenerating from scratch discards the history while producing an artifact that looks every bit as authoritative. Durability is what lets a consumer trust that yesterday's number still means the same thing today, and that when it changes, it changed because the law did.
The same lesson had already arrived at the benchmark. Chapter 8 told it: an early layer that narrated each PolicyBench result in prose was caught inventing derivations — describing elderly enrollees as beneficiaries of a "non-elderly expansion" — and any audit that read those confident narratives inherited the fabrication whole .[6] The fix grounded every explanation in the engine's actual internals and added a validator that rejects any mechanism the underlying trace does not support. Even the audit layer needs ground truth; you give the judges traces, not prose.
And once, the gate caught a lie aimed at readers of this work. When we adapted a US-calibrated population to the United Kingdom and wrote up the results — the next chapter tells that story properly — a summary statistic that no run had actually produced found its way into a draft: fluent, plausible, fabricated. A verification gate caught it before publication, and it was struck.
None of these episodes is an embarrassing exception to a spotless record. They are the record. A verification regime that only ever vindicates its builders is a marketing department with test coverage; the gates earn their standing by how often they turn red on their own authors.
The admission rule
One rule sits under everything in this chapter, and the rest of the book enforces it. Simulation is admissible only where its verification chain terminates in ground truth — a forecast score, an oracle parity, a statutory exactness, a calibrated marginal.
Each terminus is a different kind of true thing, checked a different way. An oracle parity: an independent implementation agreeing on every record, this chapter's subject. A statutory exactness: the money-atom carried to its end, an obligation traced to the quoted provision that fixes it. A calibrated marginal: an input distribution reconciled against a number the government actually counted, so the population under the rules is pinned to something real. A forecast score: a prediction published with an interval and graded only later, when the official figure lands. They are not interchangeable, and they share the one property that makes any of them admissible — each bottoms out somewhere the model's own confidence gets no vote. Where the chain exists and holds, the estimate is admissible. Where it does not, the estimate — however fluent, however cheap, however fast — is confident fiction, and the honest move is to mark the boundary rather than blur it.
Which exposes the assumption this whole chapter stands on. Everything here rests on the existence of an oracle: a UKMOD to disagree with, a EUROMOD to file bugs against, an official calculator to match to the pound. The chain terminates cleanly for the UK, for Belgium, for the US, because in those places somebody already spent years or decades building a second opinion. Most of the world's tax and benefit law is not like that. For much of it there is no reference model at all — no calculator to race, no mature implementation to disagree with, no independent answer to check. What becomes of the discipline where there is nothing to check against?
The second week of July 2026 put that question to the test.
References
- Sutherland (2013). EUROMOD: the European Union tax-benefit microsimulation model.
- Feenberg (1993). An Introduction to the TAXSIM Model.
- Richiardi (2021). UKMOD -- A New Tax-Benefit Model for the Four Nations of the UK.
- UNU-WIDER (2026). SOUTHMOD: Simulating tax and benefit policies for development.
- The Axiom Foundation (2026). Savings allowance interaction: upstream issue filed against UKMOD.
- PolicyEngine (2026). PolicyBench: Benchmarking language models on household tax and benefit calculations.