The question sounds too basic to be unanswered: what is a household's real income — after taxes, after benefits, after the two collide? A lending platform needs it to underwrite a gig worker fairly. A benefits screener needs it to tell a family what help exists. An AI assistant needs it the moment someone asks about their Child Tax Credit, unless the plan is to invent the number. Through the mid-2020s, none of them could simply buy the answer.
Call the engineer who discovers this Priya. She is a composite, drawn from fintech teams I have sat across from; her dead ends are not. Her company advances cash to gig workers, and to lend responsibly it needs each worker's true take-home pay. She tries a tax library off GitHub: federal brackets, nothing else — for her low-income users, whose finances run through credits and benefits, it misses by 30 to 40 percent. She tries a commercial tax API: priced per call, and it cannot touch benefits at all. She tries prompting a frontier model, and gets a hallucination rate no financial product can ship — chapter 8 measured exactly this. She prices an in-house build: eighteen months of engineering, obsolete the first January a legislature meets. Every path ends the same way, in a partial, fragile solution that every company facing the question rebuilds privately, again.
A market of slices
Plenty of companies sell tax software. Specialists thrive on every slice. April embeds tax filing inside other companies' apps; Column Tax has filed more than a million returns through its API and built an AI agent to help maintain its engine as the law shifts year to year .[1] Avalara computes sales tax, and was acquired for $8.4 billion in 2022 .[2] ADP and Gusto run payroll. MyFriendBen and Benefit Kitchen screen for benefits .[3][4] PolicyEngine, EUROMOD, and Tax-Calculator serve policy research. The global tax-technology market was climbing toward $60 billion [5] — balkanized into verticals, each re-encoding the same overlapping rules behind its own walls.
What had no home was the whole: income tax plus benefit eligibility plus the population-level simulation that turns one household into a national estimate. The pieces do not separate cleanly, which is the point. SNAP falls as earnings rise while the EITC climbs; the net depends on household composition, state of residence, and thresholds that interact. Because no one sold the whole, every company that needed it rebuilt the overlap privately — the same EITC phase-in, the same SNAP deductions, encoded again and again — and each private copy was a fresh chance to be wrong.
Scale has narrowed the gap — each model generation scores better on the benchmark than the last, and the book should say so plainly — but the causes underneath do not train away. Tax law changes every January, so last year's training data encodes rules that no longer apply, and a model cannot extrapolate to a bracket it has never seen. Fifty states each run their own income taxes, credits, and benefit programs — combinatorial sprawl past what pretraining memorizes reliably. Eligibility turns on dozens of interacting variables, where one wrong input silently yields a wrong result. And financial calculation tolerates neither 95 percent nor 99: right nineteen times in twenty means one return in twenty filed wrong. Column Tax's own engineers, who profit when models improve, put it flatly: language models "cannot 'do taxes' on their own because tax calculations require 100% correctness" .[1] On our benchmark, the best 2026 models still got roughly one in nine household calculations wrong by more than a dollar .[6] The way through is the tool chapter 8 specified: a deterministic, auditable engine the model calls, and this chapter builds it.
Administration is the product
Before the building comes a correction, learned the hard way, to what "good" means here. The hardest problems in benefits software are not mathematical. Whether a formula is legally correct is one question; how fast a rule change becomes publishable, how many misunderstanding loops open between the policy expert and the engineer, how much hand-holding a partner developer needs, and how often a tool hands back a misleading answer because its interface cannot express a real-world exception — these decide whether the software is any good, and none of them is a math problem. They surface as operational metrics: time to publish, discrepancy rates against authoritative examples, onboarding time, the rate at which a support ticket exposes an unwritten assumption.
Households, meanwhile, experience policy through its administration. A credit that exists in statute but takes six months to reach a screener is not fully real to the family it was written for. A calculator that silently misses an immigration rule or a state exception erodes trust, burns caseworker time, and suppresses take-up. And when the administration fails at scale, the failures have names and dates. In 2023, during the great post-pandemic Medicaid redetermination, CMS ordered states to restore coverage to roughly half a million children and families after finding a defect in the automatic-renewal logic of their eligibility systems .[7] KFF Health News documented how long fixes take inside the eligibility systems Deloitte runs for states: in Kentucky, one system limitation that cost a resident, Beverly Likens, her Medicaid coverage took about ten months, more than 3,500 hours of work, and over $500,000 to resolve; in Georgia, officials were still untangling a defect affecting more than 25,000 SNAP and TANF cases nearly two years after it was first reported .[8] Deloitte-built eligibility systems span roughly twenty-five states, under contracts that had reached some $6 billion by late 2025 .[9] Virginia Eubanks and colleagues showed the quieter version of the same failure: a screening tool can look authoritative while its logic stays opaque, so people act on wrong outputs without thinking to challenge them .[10]
Run through those cases and one trait repeats. The defect was buried in a vendor system that no outside party — often no inside party — could inspect, trace to a rule, or fix quickly. The rule was right in statute and wrong in software, and invisible until it produced a headline.
Pricing the error rate
Then policy started billing for it. The One Big Beautiful Bill Act — Public Law 119-21, enacted July 4, 2025 [11] — sharpened a decades-old lever: quality-control penalties for high error rates go back decades, but for the first time states pay a share of the benefit costs themselves. Beginning in fiscal 2028, that share is tied to the state's own payment error rate: nothing while the rate stays under 6 percent, then 5, 10, or 15 percent of benefits as it crosses 6, 8, and 10 percent — and for that first year, a state may elect whether its fiscal 2025 or 2026 error rate is the one that counts. The national payment error rate for fiscal 2025 was 10.62 percent ,[12] an average that would put many states straight into the top tier; one outside estimate put first-year state exposure at roughly $9 billion .[13] The same law raises the state share of SNAP administrative costs from 50 to 75 percent in fiscal 2027 .[11] Medicaid's new work requirements, meanwhile, oblige expansion states to stand up systems by the end of 2026 to verify that able-bodied adults are meeting an eighty-hour-a-month engagement requirement .[14] A payment error rate used to be an audit statistic. It is now a line in the state budget — and lowering it is precisely the job of rules infrastructure that is verified, current, and provenance-carrying.
The stakes of getting this wrong run past budgets, and the record is not hypothetical. Australia's Robodebt scheme raised automated debts against welfare recipients on faulty income-averaging; the courts unwound it, and a royal commission put ministers under oath — its report called the scheme "a crude and cruel mechanism, neither fair nor legal" .[15] Michigan's MiDAS system accused tens of thousands of unemployment claimants of fraud; of the determinations it issued without any human review, roughly ninety-three in a hundred were wrong .[16] Arkansas replaced a nurse's judgment with an algorithm for allocating Medicaid home-care hours, cut care for disabled residents, and drew years of litigation .[17] In each case the arithmetic was automatable and the judgment was not, and the system automated both.
American benefits law now draws that line explicitly. USDA's Food and Nutrition Service has held that AI may not replace the state merit personnel who make SNAP eligibility determinations — its framework states that "All AI must be used in compliance with program requirements for the use of merit systems personnel, such as those applicable to SNAP" , and its companion automation memo draws the same line at the job description: "bots are considered non-merit staff and prohibited from performing these merit staff functions" .[18] That boundary matches the architecture this book has been arguing for from the other direction. A human makes the determination. The tool's job is to make the determination checkable — and everything in the failures above, from Robodebt to the Kentucky renewal logic, is a case where the determination was neither human nor checkable.
The foundation
The Axiom Foundation exists to build that tool as a public good. It is a Delaware nonprofit, fiscally sponsored by the PSL Foundation and anchored by funding from Ballmer Group; Ariel Kennan became its president on July 1, 2026, and it launches publicly on July 28 .[19] Its charter is narrow and immodest at once: encode the law itself, exactly, as open infrastructure — in a form anyone can run and anyone can check. PolicyEngine had proved that tax and benefit rules could be encoded accurately at all. Axiom is the attempt to encode all of them, and the wager underneath it is the subject of the rest of this chapter: that agent-drafted encoding, behind checks that block a merge, is what makes that scope tractable for a nonprofit. A sibling institute, Thesis, takes the other half of the problem — forecasting what governments will actually do — and belongs to later chapters.
Rules that carry their sources
An Axiom encoding is not an ordinary program. Its format, RuleSpec, is a set of YAML files that hold the law's logic and the law's numbers apart — each file versioned, each testable, each stamped with the dates over which it takes effect. The logic that computes a credit lives in one place; the parameters it reads — a rate, a threshold, a phase-out amount — live in another. An act of Congress changes the logic; an agency's annual inflation adjustment changes only a parameter value. Both land as versioned files with effective dates, so an encoding answers more than "what does the law say?" It answers "what did it say last March, and when did we learn it had changed?"
Take the Earned Income Tax Credit. The statute defines a credit that climbs with earned income at a set rate up to a ceiling, plateaus, then phases out above a higher threshold that shifts with filing status and number of children. In RuleSpec that becomes a formula referencing named parameters — phase-in rate, earned-income ceiling, phase-out rate and threshold — and each parameter carries its own citation, its own effective date, its own link to the source. Watch the two kinds of change move through it. Each fall, the annual inflation adjustments arrive, and the parameter files gain new dated values while the formula — which Congress has not touched — does not move; the diff is a page of numbers with effective dates. When Congress restructures the credit, the formula changes, and the diff says so in logic rather than in numbers. A reader of the repository's history can tell, at a glance, which years the law changed and which years only its amounts did — a distinction that takes a tax lawyer to extract from the statute books.
Beneath the rules sits the corpus: the source text itself. Axiom ingests statutes, regulations, agency manuals, and sub-regulatory guidance as anchored provisions, each clause individually addressable, organized against a registry of roughly 41,000 legal concepts to verify. The corpus is what makes provenance possible — an encoding earns trust by pointing back to the exact governing words, which live in a service anyone can open and read. The anchoring matters more than it sounds, because the rule a caseworker actually applies is often buried in an agency manual or a guidance letter, not the statute; the corpus holds primary law and its sub-regulatory layers at clause resolution, so the pointer lands on the words that govern.
The gauntlet
Encoding law by hand is slow and does not scale; a team can keep one country current or attempt fifty, not both. The pipeline, axiom-encode, turns encoding from analyst-months into agent-runs. It pulls the relevant provisions from the corpus, scaffolds a prompt around them — the statute text, the concepts it touches, the shape of the encoding to produce — and hands the draft to a language model. Then, before anything merges, the draft runs a gauntlet in which every gate can block the merge: the code must compile; its tests must pass; a proof step must validate the logic's internal consistency; the output must be compared against independent reference calculators; and every monetary obligation must be grounded in its source. Only a draft that clears all of it yields a signed manifest — a cryptographic record of what was encoded, from which sources, having passed which checks — so a later reviewer, a skeptical agency, or a court can reconstruct the provenance rather than take it on faith. Human review sits deliberately at the top of this stack: the machine drafts and the gates screen, while a person adjudicates the judgment calls the statute leaves genuinely open — the cross-references and exceptions no gate can settle.
Each gate defends against a different failure. Compilation catches an encoding that does not run. The tests catch a change that breaks a case that used to work. The proof step catches logic that contradicts itself. Comparison against an independently built calculator catches the failure no self-check can see — an encoding that runs, passes its own tests, and is still wrong — which is the subject of the next chapter. And the last gate catches a number with nothing behind it.
That last gate deserves its own paragraph, because it is the pipeline's sharpest idea. The money-atom gate demands that every monetary obligation an encoding produces — every dollar figure the law commands — trace to a quoted excerpt of the governing source: the actual authorizing words, pulled from the corpus, not a section number in a comment. It targets money rather than every clause because a wrong eligibility flag is a bug, but a wrong dollar figure is a household underpaid or a program overexposed — the failure mode with a face on it. If a figure cannot be grounded that way, the build does not warn. It fails. Ordinary software cites its sources in documentation, and documentation drifts out of date the moment someone edits the code; a merge-blocking check cannot drift, because nothing merges until it passes.
Provenance stopped being a promise and became a merge condition.
Notice, finally, what the gauntlet does not trust: the model that wrote the draft. The pipeline runs several language-model backends and stays indifferent to which produced a given encoding: wherever the draft comes from, the gates hold the guarantee. A better model drafts faster and trips fewer checks. It never earns the right to skip them.
The whole of the law
By July 2026 the US repository held on the order of 3,000 rule files with matching tests, covering federal law plus twenty-eight state codes — absorbed with their full git history, so the encodings inherit years of prior fixes and the reasoning behind them, a change of representation rather than a fresh start with fresh bugs to verify. Separate monorepos carried the United Kingdom, Canada, New Zealand, and Belgium, alongside a set of African lanes validated against established reference models — the next two chapters take up the verification against reference models and the African week. A single pipeline produced all of it, which puts the marginal cost of the next jurisdiction in units of agent-time rather than institution-years.
The scope is deliberately, almost provocatively, total: not "the parts of law that compute a number" but all public policy — statutes, regulations, agency manuals, statutory guidance, grant conditions. Among the early UK encodings is the council-tax-reduction policy of the Royal Borough of Kingston upon Thames: a local scheme binding on a few tens of thousands of residents, encoded with the same machinery as federal income tax, because a resident is governed by it as concretely as by any act of Parliament.
Bindingness is metadata, never a scope filter.
Whether a provision is hard law or soft guidance is recorded as a property of the encoding; it plays no part in deciding what is worth encoding. And the encodings model authority explicitly: the chain by which a statute delegates power to an agency, and an agency binds itself through its own published policy. The doctrines are old — in US law an agency must follow its own regulations; in English law it must follow its own stated policy absent good reason citation pending — and encoding them lets a system answer not only what a rule computes but whether the body applying it had the authority to. That matters because much of law computes no amount at all; it asks whether a conclusion holds, fails, or remains undetermined, given a history of events: is the notice valid, the deadline met, the procedure followed? Modeling authority, delegation, and that second shape of question is the difference between encoding a calculator and encoding the law. Everything ships under permissive licenses — Apache 2.0 for the code, Creative Commons Attribution for the content — so no one need ask permission to run, fork, or build on the encoded law.
A commons, governed
Building this as a public good removes the sharpest objections — no shareholders, no paywall on the law — without removing the need for governance, and honesty requires naming the failure modes. First, whose rules come first: funders and large institutional users have priorities, and the risk is a roadmap that tracks whoever pays rather than where accurate rules matter most; deprioritize one jurisdiction's benefit code and you have decided which vulnerable households get accurate answers. The guardrails are transparent governance, a published roadmap, and the standing ability to fork. Second, capture: even a nonprofit can be captured by a dominant funder or a maintainer monoculture, until "open" is a brand rather than a fact; independent governance and a genuine plurality of contributors are the only durable defense. Third, staleness — the characteristic failure of public goods is not bankruptcy but quiet drift, laws changing while no one notices. The conformance ratchets of the next chapter exist to turn "keep it current" into a gate.
Openness is the mechanism behind each defense. A fintech can read exactly what the encodings compute. An agency can confirm they match statute. A citable, inspectable rule is more defensible than a model's unexplained output when a regulator or a court asks for the audit trail. And the ability to fork is the ultimate check on capture. Underneath these tactics sits a structural commitment: the encoding of law is reference infrastructure — closer to a dictionary or a map projection than to a product — and reference infrastructure is not well served by a private gatekeeper sitting between the public and the statute. So the rules layer is a commons, and the commercial activity growing around it — managed runtimes, service-level guarantees, applied products — is built on the commons, never owner of it. One such operator already exists in formation, organized around graded simulation; it is not yet public.
The foundation commits to the dividing line in public: the commercial service tier will always be a separate entity, so the foundation never competes with builders on its own commons. The law stays free to run. Only the convenience of not running it yourself has a price.
The cache that wasn't
The honest way to end a chapter about a verification regime is with the time it turned on us. In July 2026 we ran an experiment on a premise I found elegant: if every rule is deterministically derived from source text, then no encoding needs to be a durable artifact — encodings are just cache. Discard any module and regenerate it from the corpus on demand. Internally the test was called the B1 probe, and it put the premise to twenty-five modules.
It failed. Regeneration was genuinely cheap — about eight cents of compute per module — but in twenty-three of the twenty-five cases the original encoding was the correct one, and the freshly regenerated version introduced naming instability that would have silently broken every downstream consumer. Same source, same pipeline, different output. We adopted the conclusion the evidence forced, against the elegant version of our own story: an encoding is a durable artifact with provenance — stored, versioned, migrated with care — and regeneration went back on the shelf as a research problem. I count the episode as the discipline working. A project convinced of its own story would have shipped regeneration and met the breakage in production.
Which points at the gap this chapter has left open on purpose. Encoding all public policy at agent speed is worth nothing unless the encodings are right, and "a language model wrote it and the tests passed" is no reason to trust a number that decides whether a family makes rent. Compilation proves the code runs. The money-atom gate proves each dollar traces to the law. Neither proves the encoding matches the world. That proof is a separate discipline — comparison, case by case, against independent calculators built by other people with other tools, and an explanation owed for every place they disagree. How to trust law encoded by machines is the subject of the next chapter.
References
- Column Tax (2024). Will AI Agents Help File Your Taxes?.
- Business Wire (2022). Avalara to be Acquired by Vista Equity Partners for \$8.4 Billion.
- MyFriendBen (2024). MyFriendBen receives \$2.4M grant to expand access to public benefits.
- Benefit Kitchen (2026). Benefit Kitchen -- Data Analytics for a Better Life.
- Precedence Research (2025). Tax Tech Market Size to Hit USD 60.66 Billion by 2034.
- PolicyEngine (2026). PolicyBench: Benchmarking language models on household tax and benefit calculations.
- Centers for Medicare \& Medicaid Services (2023). Coverage for Half a Million Children and Families Will Be Reinstated Thanks to HHS' Swift Action.
- Liss (2024). Errors in Deloitte-Run Medicaid Systems Can Cost Millions and Take Years To Fix.
- Pradhan (2024). Medicaid for Millions in America Hinges on Deloitte-Run Systems Plagued by Errors.
- Eubanks (2020). Exposing Error in Poverty Management Technology: A Method for Auditing Government Benefits Screening Tools.
- 119th United States Congress (2025). H.R. 1, One Big Beautiful Bill Act.
- U.S. Department of Agriculture (2026). USDA Announces FY 2025 State Payment Error Rates in SNAP.
- Center on Budget (2026). States' First-Ever Bill for SNAP Benefits Could Cost Billions.
- Centers for Medicare (2026). Medicaid Program; Community Engagement Requirement for Certain Individuals.
- Holmes (2023). Report of the Royal Commission into the Robodebt Scheme.
- Michigan Supreme Court (2022). Bauserman v. Unemployment Insurance Agency.
- Arkansas Supreme Court (2017). Arkansas Department of Human Services v. Ledgerwood.
- U.S. Department of Agriculture (2024). Use of Advanced Automation in SNAP.
- Axiom Foundation (2026). Axiom Foundation -- The world's rules, encoded.