On July 8, 2026, two findings ledgers went up on GitHub a few hours apart. Ghana's listed thirteen findings .[1] Uganda's listed none .[2] By the end of that week there were five ledgers — Zambia's on the ninth, Ethiopia's after a single overnight run, Rwanda's closing the set — and behind each one stood a country whose tax and benefit system our agents had encoded, provision by provision, and raced against the reference model that UNU-WIDER maintains for it.
Those reference models deserve their billing before anything else in this chapter, because nothing in it works without them. SOUTHMOD is the family of tax-benefit models that UNU-WIDER has built and maintained with national research teams for a decade, covering fourteen developing economies [3] — GHAMOD for Ghana, UGAMOD for Uganda, MicroZAMOD, ETMOD, RWAMOD. They are the product of exactly the slow institutional work this book has described elsewhere: a national tax-benefit model has historically taken years — a team, a budget, a data-sharing agreement, a maintenance commitment measured in decades — which is why such models have lived inside finance ministries and a few well-funded institutes, and why most of the world has none at all. The fourteen SOUTHMOD countries are the exception, and the exception is what made the week possible: they were the places where a second opinion already existed. Chapter 10 ended by asking what the verification discipline becomes where there is nothing to check against. The first answer is that you begin where there is something.
Each country already had a model. What the week added was a second, independent implementation of its statutes — permissively licensed, open to anyone for any purpose, built from the gazetted law itself — plus the thing I have come to think matters more: a public ledger recording every place the two implementations disagree and why. What the reference models' maintainers gained runs the other way, and I will get to it. Everything here, as in the last chapter, is our own program reporting on itself; the ledgers are public and the comparisons are built to be rerun.
Two things had to be true at once for any of this to matter: that a model could be built this fast, and that each one arrived carrying the evidence for whether it was right. Most of this chapter is about the second condition, because without it the first is a liability.
What a finding is
Thirteen and zero, produced the same day by the same method. Neither is an obvious headline, and both are good news — but only once you know what a finding is.
A finding, in these ledgers, is a dispositioned divergence: a specific case where the encoded statute and the reference model produce different numbers, chased until someone can say why, and recorded together with the evidence for the explanation. "Explained" carries the same weight it carried in chapter 10 — an arithmetic reconciliation, a citation to the governing text, or a documented convention of the reference model, never a shrug. The one kind of result the process treats as unfinished is a divergence nobody can account for.
The explanation can land on either side. Sometimes the encoding is wrong — a misread threshold, a missing phase-out, a rate applied to the wrong base — and the fix goes into the encoding. Sometimes the reference model is out of date, or has made a defensible modeling choice the statute does not require, and the finding travels upstream: written up with its arithmetic and its citation and reported to UNU-WIDER, whose team maintains the SOUTHMOD models. Ghana's thirteen findings are thirteen such adjudications, posted where anyone can read them. The reference models are instruments, and like any instrument they are improved by being read carefully against their source. There is a quiet reversal of trust inside that sentence. The natural assumption is that the established, decade-old model is the authority and the freshly generated encoding is on trial — and most of the time that is right, and most findings are the encoding's to fix. But not all. When a divergence is chased to ground and the statute plainly says one thing while a model does another, the encoding becomes the instrument that caught it. Neither implementation is presumed correct. The statute is, and both are measured against it. Nobody grades anybody, and everybody grades the law.
Which brings me to Uganda, the result I keep returning to because of what one day used to buy. PolicyEngine UK went live in the fall of 2021 after Nikhil and I had encoded, by hand, every major tax and benefit in Britain — income tax and National Insurance, Universal Credit and its taper, child benefit and its clawback. We typed thresholds off gov.uk and argued about rounding. Uganda took one day. On July 8, the agents reused the recipe Ghana had just proven: read the gazetted acts, write the rules, race them against UGAMOD across the full test surface. The ledger recorded no findings at all — in its own phrase, the encoding was "statutorily exact everywhere tested," and the widest disagreement anywhere in the suite was four-tenths of one Ugandan shilling per year, itself examined and dispositioned as rounding: dust rather than a tolerance band .[2] I checked the arithmetic myself that evening, looking for the catch. There wasn't one. The catch was in my assumptions about how long careful work takes.
Zero findings is not the absence of a result. It is one. Uganda's empty ledger says UGAMOD turned out statutorily exact everywhere our suite reached, and that the encoding matched it to the shilling — established by evidence, not assumed from silence. A model checked against an independent reference and found to agree is in a different epistemic state from a model no one has checked, even when the two would print identical numbers. Uganda's zero is a compliment paid to a reference model, in evidence, and of the five results that week it is the one that reassures me most: the method can tell the difference between right and merely unexamined.
The rest of the week
Zambia landed on July 9, encoded in about a day and checked against MicroZAMOD: eight findings, and the first ledger with divergences on both the tax and benefit sides rather than the tax side alone — among them a definition of the pay-as-you-earn base and a set of excise rates carried a year stale .[4] Eight findings on a first pass across an entire system is low, and the both-sides detail is informative: the seam where taxes and benefits meet is where implementations drift apart most easily, and Zambia was the first lane with enough surface tested to see the drift from both directions.
Ethiopia ran in a single overnight pass against ETMOD and produced two findings, both in the same place: the minimum alternative tax introduced only months earlier by Proclamation 1395/2025 .[5] A statute that new has barely been implemented by anyone, and two first-pass disagreements about how it works is close to what you would predict — both of the understandable kind, reasonable readings of a fresh law that differ, rather than arithmetic gone wrong.
Rwanda closed the week with five findings, and one procedural first for us: the comparison ran live rather than replayed from a fixture — the encoded model and RWAMOD were handed the same eighty-six cases and asked to agree in real time. The run surfaced the five findings, and once they were dispositioned the two agreed across all eighty-six citation pending.
Tanzania is under way as I write. And Nigeria has begun — and Nigeria is different in the way this book cares about most. Nigeria is not a SOUTHMOD country. No reference model exists for it, and no administrative ground truth is available to check an encoding against. The rule of chapter 10 does not bend for enthusiasm: a simulation may be believed only where its chain of verification ends in something real — a reference model it matches, a statute it reproduces exactly, or a number the world will later publish. Nigeria's chain does not yet terminate in any of those. The encoding work can proceed, and it has; the claim that the result is correct stays weaker than the same claim for Ghana or Uganda, and the honest thing is to say so in the same breath that reports the work exists.
The boundary of the method is part of the method.
The recipe
Across five countries and five different statutes, the recipe barely changed, and it fits in a paragraph. Capture the governing acts from the official gazette, each provision anchored to a page in a signed source manifest, so every rule traces back to the exact text it came from. Encode the system one instrument at a time. Put every change through the same merge-blocking gates the previous chapters described — if it does not compile, does not pass its proofs, does not clear the oracle comparison, it does not merge. Run the oracle suites first case by case, then across a whole simulated population, so that agreement holds not just on the examples someone thought to write but on the shape of the distribution. Publish the findings ledger. Report what it turned up to the people who maintain the reference model. And once a country reaches parity, the gates ratchet: coverage can only rise and unexplained gaps can only fall, so the model cannot quietly rot back to wrong.
The cost of a country, run this way, is measured in days of agent time. For the entire prior history of the field it was measured in years of institutional effort. But the whole week was designed to make a point that cuts against celebrating the speed: a model that costs days and that no one can check is only a faster way to be confidently wrong. What traveled from Ghana to Uganda to Zambia was the discipline bolted to the speed: the gazette manifests, the merge gates, the oracle suites, the public ledger of everything that did not match. The verification shipped in the same box as the model.
The moat and the vacuum
Chapter 2 called it the analytical moat: government modelers work from confidential tax microdata with detail no public-use file carries, and outside analysts make do with files that are sampled, top-coded, and blurred. That was told as an American story, with an American statute as its cause. But the moat was never only American, and in most of the world the better word is vacuum: no gap between the government's model and a challenger's, because no public, inspectable model exists on either side. A citizen of most countries cannot check what a tax change would do to a household like theirs, because no one outside a ministry has built the tool, and often no one inside one has published it. The asymmetry Washington lives with — half a dozen institutions arguing over a tax bill within days of its introduction — is, across most of the map, closer to silence.
Consider what a public model changes around a single decision. A finance ministry proposes to broaden a VAT, or to replace a fuel subsidy with a cash grant. The question everyone cares about — who gains, who loses, by how much, across the income distribution — lives inside a model. Where the model is the ministry's alone, the answer arrives as an assertion, and the opposition, the press, civil society, and the affected household can only accept it or reject it whole. Where the model is public and checkable, the same people can run the reform themselves, disagree about assumptions in the open, and be caught when they are wrong. The ministry keeps its tool. What ends is its monopoly on being able to answer at all.
What the July week changed is the marginal cost of ending that silence. An independent, verified encoding of a country's tax-benefit law went from a multi-year undertaking to a job of days — with fourteen SOUTHMOD-family reference models standing ready to check the work, a decade of UNU-WIDER's institution-building converted into the very ground truth that makes fast encoding trustworthy. The encodings ship under Apache and Creative Commons licenses, with the findings ledgers public beside them. Openness runs in both directions here or it is not openness: the countries get an independent open implementation, and the reference models get evidence-backed quality assurance from the most exhaustive reader they have ever had.
When the data can't travel
Rules, it turns out, are the easy half. A statute is public by definition; encoding it requires the text and a way to check the result, and both have now become cheap. The harder half is the people. A tax-benefit model is only as good as its picture of who lives under the rules — at what incomes, in what households, drawing which benefits — and that picture comes from microdata, which is where openness runs out. The asymmetry is structural, and it runs opposite to intuition. A statute is written to be public: promulgated, gazetted, binding precisely because anyone can read it, so encoding it in the open takes nothing that was meant to be hidden. Microdata is the reverse — a record of actual people, protected by law and by promise, its value coming from exactly the granular detail that makes it impossible to release. Law is portable because it is already public. Microdata is stuck because it is already private.
So the summer's second experiment asked a harder question: take the calibrated synthetic population our data layer builds for the United States — populace, the commons Part II introduced — adapt it to a country it was never built for, and grade it honestly against that country's real data. The United Kingdom was the natural test case, because it has both an open reference model and a high-quality household survey to grade against. The adapted population ran through a pre-registered evaluation harness — scoring rules fixed and published before any answers were known — and was scored against held-out survey ground truth from Britain's Family Resources Survey citation pending.
The verdict was mixed, and we reported it mixed: "a credible first pass for aggregate, earnings-centric analysis — not a substitute for native microdata on benefits, distribution tails, or subnational detail." Two specifics carried the grade. One defect was found and fixed: a stale data feed had zeroed out state pensions, and the error cascaded downstream into pension-credit calculations — caught, traced to its source, corrected. One defect was found and deliberately left open: capital gains had been carried across one-to-one into a survey surface that captures almost none of them, and the honest fix is not to transplant realizations from one country to another but to model asset positions and the rules under which gains are realized — a rebuild, designed and not yet done. The transferred population is good at what surveys are good at and blind where surveys are blind, and pretending otherwise would have been easy and wrong. Part of why I trust the verdict is the episode chapter 10 already confessed: a summary statistic in an early draft of the write-up turned out to be fabricated — fluent, plausible, unsupported by any run — and a verification gate caught it before publication. The gates aim at our own output first.
Where both halves can be built, they close tightly. In Belgium, the same machinery was calibrated to twenty-one administrative targets and hit all twenty-one within 1.8 percent; on that same population, per-record social-security contributions agreed with EUROMOD to within three eurocents per person per year citation pending. Each is a small, checkable number stating exactly how close the model came, worth more than any promise that it is right.
An open referee
Belgium points at the harder case, which is most cases. For many countries the microdata cannot leave the building: national statistical offices hold survey and administrative records under legal obligations that do not bend for a research project. You cannot download the file, and you should not want to be able to. That looks like the end of open modeling in those places — you can encode their laws in the open but never open their data.
The answer we landed on is to stop trying to move the data and to move the referee instead. Give the institution that holds the microdata two things: the transferred population, and a pre-registered evaluation pack — the scoring code pinned to an exact version, the configuration frozen, and a scorecard schema that fixes, in advance, which numbers will be reported. The institution runs the pack inside its own walls, against its own protected records, and publishes only the scorecard. Because the scorecard was specified before the run, no one can fish the data afterward for a flattering statistic; the referee's questions are fixed before it ever sees the answers. No record leaves the building. The score does.
A scorecard is deliberately small — a handful of numbers, fixed in advance, saying how closely the transferred population reproduced the real one on the dimensions that matter: how the income distribution lines up, how many households the model places on each benefit, how far the tails diverge. It carries none of the underlying records and none of their risk, and it is exactly the kind of low-dimensional summary a verification chain is allowed to end in. Trust in a model of a country comes from a referee who has seen its microdata, a test written down before the answers were known, and a score published where anyone can read it. That is less than full openness, and it is enough.
Open microsimulation anywhere doesn't require open microdata anywhere — it requires an open referee.
Step back from the week and the experiments, and notice what quietly made every move legal. The rules, the data, and the prediction never had to travel together. Ghana's statutes were encoded in the open while Ghana's household records stayed home. A population built in Washington was shipped to Britain and graded there. A statistical office can run a referee no outsider is allowed to watch. Each move works only because law, population, and forecast are genuinely separable — built by different methods, checked against different ground truth, and best kept in different hands. So far in this book, that separation has been a technical convenience. The next chapter is about an organization that took it seriously enough to build itself around it.
References
- The Axiom Foundation (2026). Ghana encoding findings ledger: GHAMOD parity program.
- The Axiom Foundation (2026). Uganda encoding findings ledger: UGAMOD parity program.
- UNU-WIDER (2026). SOUTHMOD: Simulating tax and benefit policies for development.
- The Axiom Foundation (2026). Zambia encoding findings ledger: MicroZAMOD parity program.
- The Axiom Foundation (2026). Ethiopia encoding findings ledger: ETMOD parity program.