This is an Evidence Press audio briefing on The Case for Assurance Infrastructure, an essay from the Policy Identification Observatory. Cheap, capable AI agents are arriving roughly on schedule. So the question people ask is: what could a government department do with a million agent sessions? The essay argues that this is the wrong place to start. The right question is who checks the output, against what, and at what cost — because in the one deployment measured end to end, checking, not generating, was already the dominant cost, and that arithmetic only strengthens as generation gets cheaper. The technical case rests on four bounds. Redundant agents collapse: by simple statistics, a million calls that answer the same question carry the weight of one over rho independent opinions, where rho is the correlation between their errors. Serial tasks resist teams: no number of agents beats the reciprocal of a task's sequential fraction. Information is capped by the sources: mass episodes only pay when there is a vast corpus to read — one agent pass per document, which the essay calls the evidence census. And long chains of reasoning decay: at ninety-nine per cent per-step reliability, a hundred-step chain fails two times in three. Every escape from these bounds runs through verification machinery. The cost evidence says this has already happened. When the United Kingdom government used its Consult tool to categorise over fifty thousand consultation responses, the machine time cost two hundred and forty pounds — and the expert checking cost more than four times that. Assurance was over four fifths of the direct bill. And the argument holds even if the capability forecasts fail, because the same cheap generation is available to everyone who submits evidence to government. Eighteen million fabricated comments reached the American net-neutrality rulemaking before this technology even existed. The essay closes with a research programme: the avenues that would make assurance cheaper, simpler, and more capable, and sixteen specific projects, each with a falsifiable success criterion, ranked by probability of delivery when agents themselves do the mechanical share of the research. Every project earns a medal, tier by tier: gold for the near-certainties, silver for the better-than-even bets, bronze for the coin flips, and teal for the hard tail. Durable instruments — theorems, protocols, standards — outrank measurements of today's models, which fade with every model generation. Above the instruments sits a theoretical core in platinum: a compositional calculus for assured claims, the layer that credential and provenance standards leave open. With frontier models now publishing machine-checked mathematics, that theory is work agents can begin today. The gold projects are cheap, fast, and need no one's permission. The highest-payoff project in the teal tail is the least likely to happen — because only government can host it. The full essay, with the ranked table and every source, is at evidence hyphen press dot pages dot dev, under the Observatory.