Organizations restructure for ordinary reasons — a merger, a funding round, a founder's exit, a board that wants cleaner reporting lines. It is rare for one to restructure around a claim about knowledge.
This one did.
The claim had been accumulating through the previous four chapters without being stated as one. Building the engine taught us that rules, data, and prediction are separable concerns; the SOUTHMOD week and the transfer experiment showed the separation doing real work, statutes traveling openly while microdata stayed home. For most of a decade the three had lived together inside one tool — one set of repositories, one brand, one habit of mind — and the joints stayed invisible because the bundle worked and nothing pulled on them.
Verification pulled on them. Once each part of a policy estimate must prove itself against ground truth, the parts stop resembling one another. A rule proves itself against the statute it encodes and against a reference calculator — TAXSIM, UKMOD, a SOUTHMOD model — that either reproduces its answer to the last cent or forces the difference to be explained, and it does so at the moment the rule is written, before it ever touches a household. A population estimate proves itself differently: against surveys, and against the administrative totals the government already publishes — a calibration you can check the day you build it. A prediction can prove itself in neither way, because the thing it describes has not happened yet. It waits, scored only when reality arrives and the official number lands.
Three clocks
Three kinds of claim, on three clocks. A rule's truth is settled today, against a document. A population's truth is settled today, against a count. A forecast's truth is not settled until some future date when a statistical agency posts a release. And the clocks imply the failure modes. A rule fails by contradicting the law — a bug against a fixed target, findable and fixable. A calibration fails quietly, by drifting away from a population it used to resemble. A forecast fails loudly and unarguably, by being wrong about the world — a miss you cannot litigate away, only record. You do not manage those three failures with the same reflexes, or the same people, or — this chapter's claim — inside the same institution.
Run the three clocks on a single tax credit and the difference stops being abstract. Whether the credit is encoded correctly is answerable this afternoon, against the code section and an oracle. Whether the households it reaches look like the country is answerable this afternoon, against the tax agency's published totals. How many families will actually claim it next filing season is answerable only next filing season, when the agency counts them. The first two are audits. The third is a bet.
Axiom states, Thesis predicts
The reshaping took its cue from that observation rather than from the forces that usually move an org chart. PolicyEngine divided along the seam verification had exposed, into two poles that are opposite in the most elementary way two claims can be — and the names were chosen to say so. An axiom is an accepted truth: not on trial, the thing you reason from. "What does the law say?" has that shape; the answer is fixed by a statute already written, and the only live question is whether an encoding faithfully matches it. A thesis is the mirror image — a proposition advanced in order to be tested, granted no standing until the evidence rules. "What will happen?" has that shape; no document settles it, and any answer stays provisional until the world supplies the verdict. Two words, two epistemic tempers. The reorganization is an attempt to stop housing them in one body.
The Axiom Foundation, chapter 9's subject, is the determinism pole .[1] It encodes law exactly — federal statutes, state codes, agency manuals, the council-tax-reduction scheme of a single London borough — and asks its agents to state, never to guess. Its public launch is set for July 28, 2026. Everything it ships is meant to be right the way a proof is right: checkable line by line against a source, and wrong only if the source is.
The Thesis Institute is the prediction pole .[2] Its product is a public docket of forecasts about government metrics — the first print of the unemployment rate, a program's caseload, a wait time — including forecasts made under explicit policy counterfactuals. One live page asks what Medicaid call-center wait times will be in March 2027 if the work-requirement deadline chapter 9 described gets delayed: a number that does not exist yet, about a world that has not happened yet, staked out in advance. Its public launch is set for around September 15, 2026. Everything it ships is meant to be provisional — a value with an interval around it, entered into the record early enough that reality can grade it.
The two poles compose, and the Medicaid page shows how. What the rule would be if the deadline slipped is an axiom: an encoded change to the law, exact, supplied by the determinism pole. What caseloads and wait times would then do is a thesis: a forecast, owned and eventually graded by the prediction pole. A policy counterfactual is precisely that stack — an exact if run forward into an uncertain then — and separating the institutions leaves the pipeline between them intact while making plain which half of every counterfactual is settled fact and which is open forecast. One honesty the docket enforces on itself: a conditional forecast is graded on the branch the world actually takes, and the branch not taken stays a projection, labeled as one. to verify
The book's discipline thereby becomes a literal product feature. Every forecast on the docket carries a published interval and will be graded when the official number lands. As of this writing, none has resolved. The scoreboard is live; the grades arrive with reality — and there is nothing else honest to report about a forecasting institution whose first resolutions are still months away.
The engine, recomposed
What happens to the original tool? It goes to both poles, because it was always both, bundled tightly enough that nobody had to notice. For years one team shipped one release, and the rules, the data, and the projections rode out together; nothing forced the question of which was a proof and which was a guess. The oracles and the scoreboard forced it. PolicyEngine is being recomposed into the two layers that were underneath it the whole time: its rules — the thousands of encoded provisions that answer what a household owes and is owed — become part of Axiom, and the calibrated microdata beneath every population estimate becomes populace, the shared substrate both poles stand on. populace deserves one plain sentence and no more: it is the calibrated-microdata commons — households synthesized and reweighted until the totals they imply line up with the counts the government already publishes — and the entire stack rests on it, because an exactly encoded rule and an honestly scored forecast both still have to be applied to a representative population before they mean anything about a country. The model code that used to live in PolicyEngine's own repositories retires downward into those two layers. The product on top keeps its name, its interface, and its users.
That last point is where the recomposition matters to anyone outside it. The teacher checking a family's benefit cliff, the journalist scoring a bill, the congressional staffer running a distributional table — none of them sees a seam; the thing they use answers the same questions it always did. What changed is underneath. It now stands on two legs that can each be checked — a rules layer verified against statute, a data layer verified against administrative totals — rather than on one monolith whose internals you had to trust as a whole.
And the recomposition buys something the field has not had: a tally. Route the population layer into a forecast, post the forecast on the Thesis docket with an interval and a resolution date, and a policy estimate stops being a number a model asserts and becomes a claim the world will eventually grade. "This reform costs fifty billion dollars" turns into "this reform costs fifty billion dollars, here is the interval, and here is the date we will find out." The institutions that produce official estimates do grade themselves in places — CBO publishes retrospectives on its aggregate budget and economic forecasts, real self-scrutiny that Part IV will credit properly. No institution we know of keeps a standing, per-estimate ledger: a policy score is published, the vote happens, and nothing circles back to record what the world did to the number. The recomposition wires the tally in as the native output format of the prediction pole rather than a report card imposed from outside.
Wrongness, isolated
Why two institutions rather than two departments? Because the split is a firewall around one specific hazard: being wrong.
The Thesis Institute has to be wrong in public, and often. A docket that never missed would be a docket with its intervals padded uselessly wide, or one quietly declining the questions that are actually hard. The worth of a graded forecast comes precisely from the fact that some grades will be bad — visibly, on the record, with the interval printed beside the miss — because that is the only thing that makes the good grades worth anything. Public wrongness is the mission. A forecaster who cannot be caught out is not being honest, only vague.
The determinism pole cannot live that way for a moment. A calculator that a benefits screener consults, that a state agency wires into its eligibility workflow, that a tax product ships to customers, has exactly one job — to be right about what the law says — and its cardinal virtue is that it is boring: same input, same answer, matching the statute every time. Housed alongside the forecasting work, it would inherit a reputation it has no business carrying; every published miss on the scoreboard would raise the wrong question about the calculator. Technically the two share code and a population. Epistemically they make opposite kinds of claim, and the surest way to keep opposite claims from contaminating each other is to keep them under different roofs, with different promises attached to their names. One roof promises to be interesting and sometimes wrong. The other promises to be dull and always faithful.
A commons and its operators
A second design constraint predates the split. Chapter 9 argued that a layer this fundamental — the encoding of the law itself, and now the calibrated portrait of the population — is reference infrastructure, closer to a dictionary or a map projection than to a product, and is not well served by a private gatekeeper between the public and the statute. The 2026 structure gives that principle a two-sided form. On one side stand the nonprofit institutes stewarding the commons: the rules, the data, the scoreboard, all public. On the other stand commercial operators building on that commons the things production demands — managed runtimes, uptime guarantees, support contracts, applied products for the banks, agencies, and firms that need someone to answer the pager at two in the morning. That is honest work, worth paying for; a hospital system cannot run eligibility screening on a best-effort open endpoint.
The distinction turns on a single preposition. Operators build on the commons. They never own it. The foundation states the commitment plainly: the commercial service tier will always be a separate entity, so the institution stewarding law-as-code never finds itself competing with the builders who depend on it. A steward that also sells is a steward with a thumb on the scale. The operator chapter 9 mentioned is the live case — organized around graded simulation offered as a service, in formation, not yet public — and nothing about the commons depends on it. The rules and the data would stay open whether or not any company ever billed for a runtime built on them — which is the test of whether a commons is real rather than rhetorical. It must survive its operators, outlast any one of them, and owe none of them anything.
So who owns the whole? Nobody, and that is load-bearing design rather than evasion. A 501(c)(3) cannot be owned. The licenses on the rules, the data, and the scoreboard let anyone build without asking. No holding company sits above the poles, because the coherence lives in the pattern, not in any body that could be bought, captured, or pressured into changing course. Astronomy has a word for this. A constellation is official — there are eighty-eight, their boundaries chartered by the International Astronomical Union, one authority deciding where Orion ends. An asterism is the other thing: a shape everyone recognizes — the Big Dipper, the Summer Triangle — made of stars that belong to different constellations or to none, visible without any body ever declaring it. This structure is an asterism, and typography even supplies the mark: ⁂, three stars in a triangle, one for each of the three questions — law, forecast, values.
The Greeks had a word too. When the villages of Attica joined into a single polis — the act tradition credits to Theseus — they called it synoikismos: separate communities becoming one society without dissolving into one another. Tradition connects Theseus's own name to tithenai, to place, the root that also gives us thesis.
The pattern underneath
Pull back from the org chart and a repeatable pattern comes into focus, with nothing intrinsically to do with taxes or benefits. Every project in this book's second half has the same three-part shape: find a corpus that is complete, public, and — the surprising part — un-computed, sitting in plain sight in a form no one has turned into anything queryable; point agents at it; ship the verified, open version of what was there all along. Axiom encodes the world's policy corpus. populace integrates the world's microdata. Thesis forecasts the government's metrics. Three instances of a single move: the analytic stack of government, rebuilt in the open, one verifiable layer at a time.
The move comes with a gate, and the gate is this book's whole discipline compressed into a test. A corpus qualifies only if its output can be checked, per unit, against ground truth — a rule against its statute and an oracle, a household against administrative totals, a forecast against the number reality eventually prints. Where that per-unit check exists, agents can work at a scale no institution could staff, precisely because each of their millions of outputs can be caught the moment it goes wrong; the check is what converts raw generation into trustworthy production. Where the check does not exist, the same agents at the same scale produce the exact opposite of knowledge — fluent, voluminous, confidently false. That is the line separating encoding the law from generating plausible law-shaped prose, and it does not move for enthusiasm or for compute.
What changed to make the pattern viable now was the unit cost of computing the corpus. For most of history the corpus of public life sat un-computed because computing it did not pay: encoding a single statute faithfully took a trained specialist days, and there are millions of provisions across thousands of jurisdictions; forecasting the budget of every county would take analysts nobody would fund. So law stayed on paper, microdata stayed in survey files, metrics went unforecast — public in principle, computed almost nowhere. An agent can encode a country's tax-benefit system in days, hold tens of thousands of provisions in view at once, draft a forecast for every metric on a docket, and never tire. The corpus was always open. The missing piece was a way to compute it that was cheap enough to attempt at scale and disciplined enough to trust, and the gate supplies the second half.
Run the test forward and the docket of candidate corpora is long, because most of public life shares the one property that matters — already written down, not yet computed. Local zoning codes are public, largely un-encoded, and checkable line by line against the ordinance text: a national map of what may legally be built where, assembled the way the tax code was. The fiscal futures of tens of thousands of county and municipal governments are forecastable and, in time, gradable against the budgets those governments actually adopt. Bills introduced in a legislature could be scored the day they drop, run through the same engine that scores enacted law, so a proposal arrives with its distributional consequences attached. Each of these is a direction rather than a promise, and each earns its place only if it clears the same gate — per-unit verifiability against something real.
The third question
Two of the three questions now have homes. What does the law say belongs to Axiom. What will happen belongs to Thesis. Between them they cover most of what a policy argument is made of — the facts of the statute and the facts of the forecast — but not the part that decides the argument once the numbers are agreed. The oldest question is still open: what do we want?
That one has no institute, and it may be the hardest of the three to anchor in ground truth, because no statute or first-print statistic holds its truth. It sits latent in a population's own preferences — harder to observe, easier to distort, and quick to punish anyone who claims to speak for it. The embryo exists: HiveSight, a survey-simulation tool that tries to elicit and forecast what populations think the way the other layers elicit rules and metrics. The later chapters take it up and benchmark it honestly, which is the only way this book is allowed to take anything up. Rules, predictions, values: the third primitive is at once the frontier of the whole program and its least finished piece — the place where a book about computing the government becomes a book about what the government is for. There is a version of this stack whose deepest payoff is not any single number but a shared, verifiable substrate everyone can reason against, so that public disagreement can finally be about ends rather than about whose figures are right.
That is the horizon. Before values, the book owes a debt it has been running up since Part I. Every estimate in every preceding chapter arrived as a single number — a cost, a poverty rate, a marginal rate — without an honest account of its own uncertainty. This reform costs fifty billion dollars: plus or minus what, and how would we know? The prediction pole exists to answer exactly that question, to turn a bare point estimate into a graded interval with a date on it. And so, before what we want, Part IV begins with how sure we can be of what we already claim to know.