Chapter 6

The household and the society

4,746 words · 24 min · draft, July 2026


A single parent in New York earns $30,000 a year. Her child has a disability. How much does she owe in taxes, how much does she receive in benefits, and what happens if she takes an extra shift?

The answer runs through federal income tax, state income tax, payroll taxes, the Earned Income Tax Credit, the Child Tax Credit, Supplemental Security Income, SNAP, and a dozen other programs, each with its own definition of income, its own idea of who counts as the household, and its own phase-out schedule, none of them designed to fit together. No one can calculate this in their head. Most accountants would struggle. Yet the answer shapes every financial decision the family makes.

This is the problem PolicyEngine's household calculator was built to solve ,[1] and this chapter is about the two questions it turned out to answer. The first is the family's: what would this mean for me? The second belongs to anyone writing the rules: what would this mean for everyone? The same engine answers both, at two scales, and the connection between the scales is the closest thing this book has to a keystone.

What you owe, what you get

A typical low-income family might simultaneously draw on the EITC, a partially refundable Child Tax Credit, SNAP, SSI for a disabled member, Medicaid, a housing voucher, a state childcare subsidy, and free school meals — eight programs, written independently, administered by different agencies, each with its own arithmetic. The fragmentation runs through the professions that serve them: caseworkers specialize in one program, tax preparers work one side of the ledger, benefits counselors know eligibility but not tax consequences. It is why eligible families leave benefits unclaimed, and why a raise can quietly cost more than it pays.

The calculator's promise is to put the whole ledger on one screen. It walks through household composition — adults, children, ages — then income by source: wages, self-employment, investment, retirement. Then the specifics that drive eligibility, the questions a tax form never asks but a benefit office always does: housing costs, childcare expenses, disability. From those answers it computes dozens of outputs under current law — each tax owed, each benefit due — ending in a single net figure for what the family keeps .[2] The US model spans federal income and payroll taxes, the EITC and CTC, SNAP, SSI, WIC, TANF, and dozens of state programs; the UK model spans income tax, National Insurance, Universal Credit, Child Benefit, and the full run of means-tested support. Each program is a separate module, drawn from statute, regulation, and agency guidance.

The interactions force that breadth. A family's SNAP benefit depends on net income after taxes and deductions; its taxes depend on credits tied to the same children who might also generate SSI or a childcare subsidy. Change one input and the effects ripple through programs that have never heard of each other. The only way to capture the ripple is to evaluate every program at once, for the same household — which no single-program calculator, however good, can do.

Rates above a billionaire's

Plot the New York family's net income against its earnings and the line has a hidden structure. There are flat stretches, where more work barely raises take-home pay. There are vertical drops — cliffs, where crossing a threshold loses a benefit worth more than the raise .[3] Chapter 4 met this arithmetic in a single family — my brother's cliff, where earning more would have cost him the attendant care he lives by — and the calculator's first job is to make that same structure visible to any family before it stumbles over it.

The mechanics stack. As the parent's earnings climb from $20,000 toward $30,000, SNAP withdraws about thirty cents of each additional dollar of net income. The EITC phases out at 15.98 cents on the dollar for a family with one child, 21.06 for two or more .[4] SSI takes back roughly fifty cents per dollar of countable income. Income and payroll taxes begin to bite. Stack them and the combined marginal rate can approach eighty percent — more than double what a hedge-fund manager pays on his last dollar. A low-income worker often faces a higher marginal rate than any billionaire, because benefit phase-outs pile on top of taxes; and this is the ordinary structure of a low-income budget, not an edge case. It also disappears into any average: blend this parent's marginal rate with her neighbors' and the eighty-percent stretch dissolves into something unremarkable, visible only at the resolution of one household — which is why chapter 1's argument for computing household by household was never really about computational taste. The same stacking happens wherever means tests overlap: in Britain, Universal Credit's taper interacts with income tax, National Insurance, and council-tax support to push some working families' marginal rates above 70 percent.

At a true cliff the marginal-rate chart spikes past 100 percent — earn more, keep less — and the calculator shades the earnings range where extra work barely moves net income: a dead zone, drawn on the family's own chart. Whether the New York parent's extra shift pays at all depends on where in that structure she stands, which is exactly what no pay stub, no program manual, and no national average will tell her. That visibility serves two audiences at once. A family can plan around a cliff it can finally see. And a policymaker can see work disincentives that no single program creates, because each cliff is an interaction — the fault of the combination, which means the fault of nobody, which is how it survived.

Three families

Three worked cases show what the calculator catches that intuition does not. Each is an interaction between rules that are defensible alone, which is why no program's own manual warns about any of them.

The first is a cliff created by a courtesy. Receiving SSI confers categorical eligibility for SNAP — qualify for one and you qualify for the other, regardless of SNAP's usual income test. As a disabled worker's earnings rise, SSI tapers gently toward zero; the day it reaches zero, categorical eligibility vanishes, and the household must suddenly pass SNAP's own income test, which it may fail outright. The gentle slope ends in a drop, positioned exactly where someone is working their way off the program.

The second is the marriage penalty. Take two single parents, each earning $25,000, each with one child. Separately, each qualifies for a substantial EITC and Child Tax Credit. Marry them, and a combined $50,000 income pushes into phase-out ranges that two $25,000 incomes never touched; the merged household can end up with less than the two households had apart. No line of the tax code says "penalize marriage." The penalty emerges from programs' differing treatment of household structure, and the calculator surfaces it the only way it can be surfaced: run the two adults separately, run them as one household, and compare.

The third arrives on a birthday. A family receiving SSI for a disabled child hits a discontinuity when the child turns eighteen: the program stops counting the parents' income and starts counting the child's own. The SSI payment itself often rises. But the same birthday can move the child out of pediatric Medicaid provisions, school meals, and childcare subsidies — family-based benefits keyed to age, not need. Nothing about the child's disability changed overnight. The rules did. For the New York parent from the opening of this chapter, that third case is a date on her calendar, and the calculator can show her the whole discontinuity years before it lands.

The what-if machine

Everything so far describes current law. The calculator's second mode edits it. Open the policy editor and the parameters of the system become movable: raise the EITC, smooth the SNAP cliff, add a child allowance, make the Child Tax Credit fully refundable, replace the lot with a basic income, flatten the income tax. The current-law world draws in gray, the reform in blue, and the charts update live for the family on the screen: two net-income lines, two sets of dead zones, stacked marginal-rate curves showing where the reform flattens the benefit system — and where it carves a new cliff somebody has not noticed yet.

Run it on recent history. In 2021, Congress temporarily raised the Child Tax Credit to $3,600 for children under six and $3,000 for children six to seventeen, and made it fully refundable. For a single parent earning $15,000 with two young children, the engine computes roughly $1,800 under the old rules — held down by a phase-in that only credited earnings above $2,500 — against $7,200 under the expansion. When the expansion lapsed at the end of 2021, that family's income fell by more than $5,000. Today the same editor measures any proposal against the $2,200-per-child credit that 2025's One Big Beautiful Bill Act set .[5] A legislative aide who once proposed a reform and waited weeks for a score can model dozens of variants in an afternoon — the benefit at $3,000 or $4,000, the phase-in from the first dollar or from $2,500, the cap at $75,000 or $150,000 — and read a different household chart for each.

It's your life, under different rules.

Estimates, not determinations

Every household calculation ships with a caveat, displayed rather than buried: PolicyEngine provides estimates, not benefit determinations .[1] The model encodes the rules as written; an actual award turns on caseworker discretion, documentation, asset tests, and local variation that no rules engine sees. The team checks continuously against official calculators and published tables [6] and treats every discrepancy as a bug to investigate — while admitting that perfect accuracy in a system this complex is unattainable, and that claiming it would be worse than naming the gap.

Within those limits, the calculation became a building block for other people's tools. The household view had been born in the UK — the 2021 launch chapter 5 described [7] — and it matters most exactly where the system is most tangled, which is why the applications found it. Build the computation once and let everything call it: the API argument chapter 5 made, vindicated by who showed up. In 2025, MyFriendBen — a benefits-screening service — launched in Colorado, North Carolina, New York, and Illinois on PolicyEngine's API :[8] plain-language questions in a dozen languages, and in about six minutes a household learns which programs it likely qualifies for. In Colorado, users surfaced an average of $1,500 a month in benefits they appeared eligible for but had not claimed .[9] PolicyEngine's own note reported that the estimates matched expected amounts more than 90 percent of the time — a self-reported figure, not an independent audit, and I flag it as such .[8] A newer prototype reorganizes the whole calculation around life events — a birth, a move, a marriage, a job loss, turning 65 — because that is how people actually meet policy .[10]

The household view, it turned out, was a primitive — and in two senses at once. The obvious sense: benefits screening, financial coaching, guaranteed-income design, and legislative analysis all separately need the same basic computation, so build it once and let everything call it. The deeper sense took longer to see. The household calculation is the atomic unit of ground-truth verification — the level at which an encoded rule can be proven against a reference calculator, an independent implementation of the same statute, or against what a caseworker actually determined. There is no such test for a national total; there is for one family's SNAP benefit.

A national estimate is nothing but a weighted sum of household calculations. If the household calculation is wrong, everything built on top of it is wrong in ways no aggregate will reveal.

That is what separates microsimulation from macro summary. A macroeconomic model might say a tax cut costs $100 billion and lifts GDP by 0.3 percent. Microsimulation can say that and also tell you the same cut hands $12,000 to a high-income household in Connecticut and $200 to a low-income household in Mississippi — and that the Mississippi family now faces a higher marginal rate, because the cut nudged it into a phase-out. The rest of this chapter is about the summing: how one family's arithmetic becomes a country's.

From one household to a country

Distributional analysis has a founding text. In 1974, Joseph Pechman and Benjamin Okner of Brookings published Who Bears the Tax Burden?, the first comprehensive attempt to allocate US tax burdens across the income distribution using household-level data .[11] They merged records for 72,000 households and found the tax system roughly proportional across most of the distribution — a result that irritated everyone, challenging conservatives who called the system punishingly progressive and liberals who called it a regressive sham. Annoying both sides at once is roughly what you would expect from the first method that forced the argument through household records instead of anecdotes. Their method — run micro-records through explicit incidence assumptions, then sum — became the template for the distributional tables that Treasury's Office of Tax Analysis, the Joint Committee on Taxation, and the CBO produce to this day.

The method's first constraint is that you cannot survey everybody. The monthly Current Population Survey samples about 60,000 households ;[12] its annual income supplement reaches seventy-five to ninety thousand — either way, a sliver of some 130 million American households. Each surveyed household therefore carries a weight saying how many others it stands in for: a rural Wyoming household might represent 5,000 similar ones, a Manhattan household 500, weights the Census Bureau computes to correct for sampling design and non-response. To score a reform, the engine computes the change for every sampled household, multiplies each by its weight, and sums. Simple arithmetic — and an inheritance mechanism, because whatever is wrong with the records passes straight through the weights into every national number built on them.

The data underneath

Plenty is wrong with the records. The survey does not merely blur the picture; it biases it. High incomes are top-coded and under-reported. Benefits are undercounted — Bruce Meyer's linkage studies, which chapter 3 introduced, found 40 to 50 percent of SNAP recipients tell the survey they receive nothing, with the same pattern across programs .[13] The sample runs too thin for most state-level work, and the asset questions are sparse. Feed that survey to a flawless rules engine and the biases choose a direction: you understate the cost of reforming means-tested programs, because the data is missing benefits, and you understate the revenue from taxing top earners, because the data is missing their income. The errors land exactly where policy debates live.

The fix is to calibrate the survey to administrative reality — adjust it until its totals match what the tax and benefit agencies actually recorded. PolicyEngine's milestone on that road was the Enhanced CPS, released in August 2025: five datasets integrated, reported taxes and benefits replaced with computed amounts for internal consistency, income distributions corrected against IRS records, and the survey weights re-solved to match 9,168 administrative totals — cutting deviations from those targets by about 97 percent to verify. Matching the targets you optimized against is fit, not proof — the honest reading of that 97 percent, and the reason the next chapter pauses on what calibration cannot guarantee.

That data work has since been rebuilt as populace: an open commons of calibrated microdata, run the way software infrastructure is run rather than as the private appendix to somebody's tax model. Where the survey is thin, it synthesizes records — a synthetic population, artificial households assembled from public data and calibrated so the whole behaves like the country without exposing any real person's record. Where variables are missing, it fills them in — imputation, predicting what the survey never asked, a method the next chapter takes apart. It calibrates weights to administrative aggregates, and it publishes certified releases the way software ships versioned builds, so an analysis names the exact population it ran on. One consequence reshapes local analysis: a city or state estimate becomes a filter on one national calibrated dataset, not the fifty-first bespoke dataset built by the fifty-first team.

Scoring a reform

With rules encoded and a population in hand, scoring a reform is two runs of the same engine: simulate every household under current law, simulate again under the reform, and sum the weighted differences in tax revenue and benefit spending .[14] Each of the roughly 100,000 records in the enhanced dataset — more records than the survey it started from — triggers thousands of eligibility checks and bracket calculations, and the whole run finishes in seconds. That speed is bought up front: building the engine took years; scoring the marginal reform takes hours because the fixed cost is already paid. Compare the traditional pipeline — read the legislation, code it up, run the model, write the memo, clear review — which takes weeks or months, by which time the legislative moment has often passed. Speed changes what analysis is for, and so does form. A think-tank report is an argument about whether to trust the authors. Infrastructure anyone can run moves the argument to whether the model is sound, and that fight can actually be settled: model quality is testable in ways a report's credibility is not.

A score is also a bundle of choices that a single headline number hides. Scores are static by default — behavior held fixed, the reform's day-one effect. Layer in labor-supply elasticities and the estimate moves: assume people work more in response to a reform and the cost comes down, assume they work less and it goes up, which is one reason two honest analysts publish different numbers for the same bill. The scoring year matters. And most subtly, every score is measured against a baseline: the world the reform is compared to, which is itself an assumption about what current law would become. Before July 2025, scoring an extension of the 2017 tax cuts required first deciding whether the baseline assumed those provisions would expire on schedule or continue — a choice worth hundreds of billions of dollars before the reform itself was modeled. OBBBA settled that particular question by making the provisions permanent ,[5] but the general lesson stands: a cost estimate without its baseline is a number without a denominator. And a cost figure without its distribution is barely a number at all — fifty billion dollars concentrated on families in poverty is a different policy from fifty billion spread evenly across the population, and only the distributional table tells you which one a bill is.

Here is what a score looks like when I run one myself. Take the Child Tax Credit under current law and remove its refundability cap and its earnings phase-in, so that families who owe no federal income tax receive the whole credit. Scored against 2026 law, my own static calculation puts the net cost at roughly $23 billion for 2026, with child poverty falling by about 2.6 percentage points — on the order of 1.9 million children. That is an author-run model output, not an independently verified fact, and it carries every caveat this chapter has raised. Its distributional shape, though, follows from the mechanics: the gains go to families whose credit currently exceeds the income tax they owe — low- and moderate-income working families with children, who forfeit part of the credit today — while higher-income households, already claiming the full credit, see nothing change.

Poverty, and what to show

Scores need outcome measures, and the most consequential is poverty. PolicyEngine uses the Supplemental Poverty Measure, which — unlike the older official measure — counts taxes and in-kind benefits like SNAP, adjusts for geographic differences in housing costs, and nets out work and medical expenses; for two adults and two children renting, the 2024 SPM threshold is about $39,400 .[15] The choice of measure decides what the model can see: a poverty measure that ignores SNAP cannot register a SNAP reform; the SPM can, which is what makes it the right yardstick for a model whose whole subject is taxes and transfers. In a 2023 analysis, PolicyEngine put roughly 9.6 percent of Americans below their threshold, with children poorer than working-age adults and seniors, and women poorer than men [16] <span class="mark mark--verify" title="date this estimate's vintage and reconcile with the Census official SPM of ~12.4% for 2022 .[17] Splitting the microdata exposes what a single rate conceals: WIC, a program that looks gender-neutral on its face, lowers overall poverty by about 0.8 percent but child poverty by 2.6 percent, deep poverty by 2.2 percent, and women's poverty by 0.9 percent against men's 0.7 [16] — the program lands where the mothers and children are.

The same runs slice by income decile, by wealth decile — using wealth imputed onto the survey — by age, by sex, by geography, and by race where the data supports it .[14] Even inequality comes in flavors: the Gini coefficient responds most to the middle of the distribution, so a reform transforming the very bottom can barely move it, while the Atkinson index, tuned by an aversion parameter, weights the bottom heavily and moves sharply. The two can disagree about the same reform, instructively.

Which raises the question of what the tool should say about any of it. PolicyEngine reports the outcomes — budget, poverty, inequality, the distribution across every cut — and adds no editorial verdict. It documents its behavioral assumptions and data limits and keeps the methodology open to challenge. I will not pretend that makes it neutral: choosing to display a poverty rate implies poverty matters, and framing results by income decile frames the debate a particular way. The honest claim is narrower — a deliberate separation of analytical infrastructure from advocacy, so that people who disagree about what to value can at least argue from the same arithmetic.

Solid and soft

Not every number the engine produces deserves the same confidence, and the tiers are worth stating plainly.

Household calculations are the most reliable output. They depend on the encoded rules, not on the data distribution: for a family with stated characteristics, the answer is exactly as good as the statute encoding, no better and no worse for whatever the survey missed. Directional and comparative results come next — whether a reform helps low-income families more than high-income ones, whether Reform A costs more than Reform B — because distribution shapes and reform ratios hold steadier than absolute levels, and the data's limits press on both sides of a comparison. Absolute budget scores deserve the most caution: PolicyEngine's aggregate revenue estimates have historically run well below official totals — by roughly a third to verify — because survey microdata undercounts top incomes and under-reports benefits, and a static model omits behavioral and macroeconomic feedback. For an order-of-magnitude figure or a cross-reform comparison, that is workable; for a precise revenue estimate, the Joint Committee on Taxation and the CBO remain the standard, because they compute on confidential return data with the top of the distribution intact.

Calibration narrows the gap and is itself checked — hold out some administrative targets, calibrate to the rest, and test the reweighted sample against what it never saw. But a 97 percent improvement on a large gap still leaves a gap, and the tiers above are where it lives. What separates an open model from a closed one was never the absence of limits. An open model prints its limits where anyone can read them; a closed model's limits are whatever institutional authority says they are.

Two reforms, seen clearly

Two episodes show what the society view catches while a debate is still live — and where its confidence should stop.

The first is the 2021 Child Tax Credit expansion. The American Rescue Plan raised the credit and, more consequentially, made it fully refundable — removing the earnings phase-in that had excluded the poorest families, the same phase-in the what-if machine priced for one family above — and paid half of it monthly from July through December. Before enactment, Columbia University's Center on Poverty and Social Policy projected the broader relief package could cut child poverty by more than half .[18] Then the Census data arrived: SPM child poverty fell from 9.7 percent in 2020 to 5.2 percent in 2021, a record low .[19] The attribution demands precision, because the debate routinely conflates two numbers the microdata keeps distinct: the credit as a whole kept 2.9 million children out of poverty that year, and the expansion alone accounted for 2.1 million of them .[20][21] The one-year cost ran to roughly $105 billion citation pending. When the expansion expired in January 2022, Columbia's researchers tracked child poverty nearly doubling within months. The forecasts landed for a reason worth naming: the expansion worked through mechanical channels — direct cash, on simple eligibility rules — where the model's arithmetic is nearly the whole story and behavioral uncertainty barely enters. That is microsimulation's best case; high-end tax changes, which turn on behavioral assumptions the data cannot pin down, are its worst. And one subtlety belongs in the record: because the expansion also raised the maximum credit for every eligible family, the distributional tables showed the largest percentage gains at the bottom but real dollar gains well up the distribution — a broad coalition or loose targeting, depending on values no model can adjudicate.

The second episode is British, and it is a diptych. One panel is the September 2022 mini-budget — the same-day analysis chapter 5 recounted. The society view is what gave that analysis its content: beyond the decile table, the engine put the package's inequality effect in numbers, a 1.2 percent rise in the Gini coefficient and a 2.2 percent rise in the top-1-percent income share .[22]

The other panel is the Universal Credit taper cut from the Autumn 2021 Budget: Chancellor Rishi Sunak cut the rate at which UC is withdrawn as earnings rise from 63 to 55 percent and raised the work allowance by £500 — a reform costing about £2.8 billion a year, cutting poverty by around 3.1 percent, with the gains flowing to roughly 12 percent of the population .[23] For a single parent working 25 hours a week at minimum wage, the change meant keeping about £1,000 more a year citation pending, while a two-earner couple gained less.

Set the taper cut against the pandemic's £20-a-week UC uplift and the shapes invert: the uplift reached everyone on Universal Credit while the taper cut reached only working recipients, so per pound the uplift cut poverty roughly 40 percent more effectively citation pending — while the taper cut did what the uplift could not, dropping the share of workers facing marginal rates above 70 percent from 26 to 9 percent citation pending. Similar money, the same program, opposite signatures: one buys poverty reduction, the other buys work incentives, and the society view is what lets a country see the trade before choosing.

By 2026 the machinery had started watching legislation as it moved — flagging which bills the model could already score and routing the rest toward encoding ,[24] the on-ramp to a much larger encoding project that Part III describes. But every capability in this chapter, household and society alike, is one engine running one computation at two scales, and it holds up only because three things beneath it hold: the rules encoded correctly, the data actually representing the people, and the behavioral responses — the seam where the model stops describing law and starts predicting humans — handled honestly. Those are the three ingredients of any microsimulation. The next chapter takes them one at a time.

References

  1. PolicyEngine (2022). Estimating Your Supplemental Security Income Benefits in PolicyEngine.
  2. Ghenis (2022). PolicyEngine's 2022 Year in Review.
  3. Ghenis (2023). How Would Reforms Affect Cliffs?.
  4. Tax Policy Center (2024). {EITC.
  5. 119th United States Congress (2025). H.R. 1, One Big Beautiful Bill Act.
  6. PolicyEngine (2024). PolicyEngine UK Model Validation.
  7. UBI Center (2021). Introducing PolicyEngine UK.
  8. PolicyEngine (2025). MyFriendBen Launches in North Carolina, Using PolicyEngine API.
  9. MyFriendBen (2025). New Tool \& Partnership Helps Pueblo Families Save \$1,500/Month.
  10. PolicyEngine (2026). Crossroads: Life-event Tax and Benefit Simulation.
  11. Pechman (1974). Who Bears the Tax Burden?.
  12. U.S. Census Bureau (2024). Current Population Survey (CPS).
  13. Meyer (2015). Household Surveys in Crisis.
  14. Woodruff (2023). From Idea to Impact: Scoring a Policy Reform on the New PolicyEngine US.
  15. U.S. Bureau of Labor Statistics (2024). 2024 Research Supplemental Poverty Measure Thresholds.
  16. PolicyEngine (2023). Breaking Down US Poverty Impacts by Sex.
  17. U.S. Census Bureau (2023). Poverty Measure That Includes Government Assistance Increased to 12.4\% in 2022, When Pandemic Relief Ended.
  18. Parolin (2022). Sixth Child Tax Credit Payment Kept 3.7 Million Children Out of Poverty in December.
  19. Fox (2022). The Supplemental Poverty Measure: 2021.
  20. U.S. Census Bureau (2022). The Supplemental Poverty Measure: 2021.
  21. U.S. Census Bureau (2022). The Impact of the 2021 Expanded Child Tax Credit on Child Poverty.
  22. Ghenis (2022). Tax cuts in Prime Minister Truss's Growth Plan 2022.
  23. Woodruff (2021). Analysing Autumn Budget Universal Credit reforms with PolicyEngine.
  24. PolicyEngine (2026). State Legislative Tracker.

Society in silico · draft in public · source