The Dutch childcare-benefits scandal drew 6.45 million euros in fines and has passed 7.2 billion to repair. The review that was meant to catch it ran thousands of times and affirmed every decision, because each one was correct.
A parent in the Netherlands underpays their share of a childcare bill by a small sum, or files a contract a few weeks late, or leaves one signature off one form. Months later a letter arrives from the Tax Administration. It does not ask for the small sum. It sets the entire year's childcare benefit to zero and demands all of it back, tens of thousands of euros, at once. The parent appeals. A court reviews the file, confirms that the paperwork was indeed incomplete, and upholds the demand. Everything in that sequence is correct. The parent is still ruined.
Run that sequence about twenty-six thousand times, roughly 26,000 being the count the parliamentary inquiry carried for 2005 to 2019 (a broader figure of at least 35,000 is a different count, not a correction of it), and you have the toeslagenaffaire, the Dutch childcare-benefits scandal: the affair that brought down a government, produced the largest fines the country's data-protection regulator had ever issued, and is still generating fresh repayment demands to the same families in 2026. It is usually told as a story about a racist algorithm. That version is true and it is adjudicated, so I am going to spend the words elsewhere. The more useful story, especially if you build systems that make decisions at volume, is an accounting story. Trust in an administrative system is a balance. A risk score can overdraw it silently, for a decade, while every single withdrawal clears.
The load-bearing fact of this whole affair is not the algorithm. It is a rule with a name: alles of niets, all-or-nothing.
Under the rule as it stood, if a parent could not fully document their required personal contribution to childcare costs, the Tax Administration was obliged to set the entitlement to zero and reclaim the entire benefit, not the shortfall. A co-payment short by a little, a late contract, a missing page: the same maximal outcome every time. And here is the part that matters. That was not a rogue caseworker exceeding their authority. It was the law as the highest administrative court read it. The Administrative Jurisdiction Division of the Raad van State, the Council of State, established the all-or-nothing line in 2011 and 2012 and did not abandon it until 2019. In a remarkable self-reflection published in 2021, the Division apologized: it had held too long that full recovery was legally required even for very small defects, and it should have offered these parents better legal protection.
Sit with what that means for anyone who has ever relied on a review step to catch a bad outcome. The check that was supposed to catch the aggregate harm was judicial review, and judicial review was not absent. It ran. It ran thousands of times. And it affirmed each decision, because each decision was correct under the rule the court itself had written.
An appeal is a per-file instrument. Point it at one case and it answers accurately: was this benefit correctly reclaimed given this file? Yes. It has no question inside it that can be asked about twenty-six thousand files at once. There is no place on the form for "is a rule that produces this result at this scale defensible?" The instrument was working perfectly and measuring the wrong unit.
If you write software, you have met this failure. It is the codebase where every unit test passes and production is on fire. It is the distributed store where each node is internally consistent while the invariant that spans the nodes quietly breaks. It is the dashboard that is green per request while the thing users actually experience, the aggregate over a day, degrades past the point of usefulness. A system can be locally correct at every point and globally indefensible, and if your only instrument asks the local question, you will not merely miss the global failure. You will accumulate a paper trail of affirmations that proves, case by case, that nothing went wrong. The overdraft is not drawn by the fraud model. It is drawn by a correct rule applied correctly, at volume, by everyone, including the institution whose job is to say no.
The algorithm did its own damage, and by 2021 that damage had a price tag written by the state itself. The Dutch data-protection authority, the Autoriteit Persoonsgegevens, fined the Minister of Finance 2.75 million euros in a decision dated 25 November 2021 and announced on 7 December. What makes the fine worth reading is that it is itemized, and the itemization shows that "the nationality field" was really three separate practices.
The regulator charged 750,000 euros for retaining applicants' dual nationality with no lawful purpose, a practice that touched roughly 1.4 million citizens between 2014 and 2020. It charged a million euros for using nationality as an indicator in the risk-classification model itself, from March 2016 to October 2018. And it charged another million for querying nationality in fraud detection between 2013 and early 2019, including targeted sweeps: 6,074 Ghanaian applicants in one, 363 Bulgarian applicants in another. The regulator deliberately escalated the second and third charges to its highest penalty band, citing the discriminatory character of the processing, the size and dependent position of the affected population, and false statements that had obstructed its investigation.
A second fine followed on 7 April 2022, 3.7 million euros, then the largest the authority had ever levied, for a fraud-signaling blacklist called FSV that held some 270,000 people, recorded characteristics including ethnicity and origin, rested on no lawful basis, and in several cases simply held data that was wrong. Together the two fines come to 6.45 million euros. Hold that number.
I am stating the discrimination as a fact rather than arguing it, because a parliamentary inquiry and a regulator already did the arguing. The inquiry report is Ongekend onrecht, a title English-language coverage quotes as "Unprecedented Injustice." It landed on 17 December 2020 with a finding that was constitutional rather than technical: the parents did not receive the protection they were owed, and group penalties violated fundamental principles of the rule of law. The third Rutte cabinet resigned on 15 January 2021. And underneath the fines and the resignations sits a number that appears on no balance sheet and has no compensation line: more than two thousand children were removed from their homes in families caught by the scandal.
Here is why 6.45 million is the wrong number to remember, and why I asked you to hold it anyway.
The operation to repair the scandal was first budgeted at around 310 million euros in 2020 and 2021. By 2024 reporting the tally had passed 7.2 billion, with a worst-case projection up to 14 billion and 9.3 billion reserved for the compensation scheme alone. The average payout per victim under the revised method had climbed to roughly 128,000 euros, well above the 30,000-per-parent figure the scheme opened with. Set the penalty against the repair and the ratio is stark: the regulatory fine is about one one-thousandth of what the failure costs to fix.
That ratio is the whole thesis in a fraction. If you thought the fines were the price of the overdraft, look again. The state paid a token of interest and is still servicing the principal six years on. The fine is a price signal, and a useful one, but it is paid out of the same treasury that is paying the billions, which makes it closer to moving money between two of your own accounts than to a penalty that changes behavior.
And the repair has reproduced the original defect in mirror image. The scandal was a machine that demanded exhaustive documentation from people who could not produce it and defaulted to the harshest outcome. The compensation scheme is a machine that demands exhaustive documentation from the same people and defaults to delay, a "light assessment" followed by a "full assessment," precise and slow enough that a Dutch university's public explainer is titled, plainly, why the settlements are so slow. De Correspondent has described a compensation affair forming inside the benefits affair. None of this is history. In February 2026 families in the scandal were reported as shocked by fresh repayment demands, and in March 2026 the new cabinet was reported brushing off the children's ombudsmen again. Write it in the present tense or you will get it wrong.
The brief that sent me here asked the right question: after a failure like this, which reforms restore the actual control, meaning a human who can see the whole model and stop it, and which restore only the appearance? The Dutch record answers with unusual clarity, because the reforms are far enough along to audit.
The visible reform is the Algorithm Register, a public catalog of government algorithms launched on 21 December 2022, meant to make their use "accessible and understandable." Publication is the reform everyone can point to. It is genuinely worth having. It is also not a brake, and the register itself demonstrates why in a single number. As of September 2026 it holds about 1,542 algorithm descriptions, up from its thousandth entry in July 2025, against a ministry target of 1,600. The obvious analysis is coverage: what fraction of government algorithms are registered? You cannot compute it. The total number of algorithms in use across Dutch government is not known to anyone, including the ministry that runs the register. At the municipal level someone did check, and found that 18 of 47 councils said they use algorithmic systems while having no entries in the register at all.
A register reports what was entered into it. Absence from it is not evidence of absence in the world, and a rising count is not rising coverage when nobody holds the denominator. That is precisely the error that started the scandal: an instrument's output, in that case a risk score, read as a fact about the world. The 1,542 is the register's reading of itself, not the government's algorithm count, and if you let it stand in for coverage you are repeating the original mistake in a nicer font.
The reform that actually produces control is duller and it works. The Netherlands Court of Audit, the Algemene Rekenkamer, keeps auditing government algorithms and keeps publishing what it finds. In 2022 it examined nine and reported that three met the basic requirements and six did not. In 2024 it looked across 69 organizations and found most public bodies using AI without a clear understanding of the risks. And in its 2024 annual audit it examined three risk-prediction algorithms, one of which delivers the finding this entire essay is built to reach.
Dienst Toeslagen, the benefits service rebuilt after the scandal, now runs a risk-prediction algorithm to find parents who need help repaying childcare benefit. It works: nearly 8,000 parents have been offered that help, and 80 to 90 percent of them were surfaced by the algorithm. The auditors also found that sensitive data about this vulnerable group is not protected, because the algorithm was not compliant with the GDPR. Read it slowly. The organization created to repair a scandal caused by unlawful profiling of a vulnerable group is itself profiling that same group unlawfully, this time to help them. Benign intent, same defect. Publication did not prevent it. The algorithm is known, named, registered, aimed at a kind outcome, and still processing sensitive data about vulnerable people outside the law. What caught it was not the register and not an appeal. It was an auditor with a statutory mandate who could look at the whole system and write, in public, "this does not comply."
That is the closest thing in the record to the human who can see the whole model and stop it, and it is worth naming its one weakness rather than papering over it: the auditor's output is a report, not a stop button. It can see the portfolio, which the appeals court could not, and it can say so where everyone can read it, which the register's own count cannot. But it does not halt the system. The brake is still missing. What exists is an eye, wired to a printer, not to a switch. The most underrated reform in the whole story is quieter still and lives in the law: abandoning the all-or-nothing rule and letting proportionality back into administrative decisions removed the rule that made each harsh outcome individually correct. That is the fix that drains the overdraft at its source, because it changes what "defensible per file" means.
You are not running a national tax administration, but you are almost certainly running something that decides at volume and gets reviewed one case at a time, and the failure mode transfers exactly.
First, ask what unit your review instruments measure. If everything that reviews your decisions, the appeal, the exception queue, the on-call runbook, the manager sign-off, asks only "is this individual case correct?", then you have no instrument for "is the aggregate defensible?", and the two can diverge with no individual error anywhere. Build the portfolio question in on purpose. Sample outcomes across a protected group, chart how often the harshest outcome fires, watch the ratio and not just the instances. The all-or-nothing rule would have been visible in ten minutes to anyone who plotted the share of claimants hit with a full clawback over time. Nobody owned that plot.
Second, treat every count as a reading until you can name its denominator. A dashboard that shows 1,542 of something is telling you what it ingested, not what exists, and a number that only goes up is not progress if the thing it is a fraction of is unknown and possibly growing faster. Before you report a count as coverage, write down what it is a count of and what it is not.
Third, wire the eye to a switch. Visibility, audits, and registers are all necessary and none of them is a brake. The Dutch record shows an auditor who could see the whole model, say it was unlawful, and watch it keep running, because saying so was the end of that person's authority. Decide before you ship who can see the aggregate and whether that same person or process can halt it. If the one who can see it can only file a report, you have built the appearance of control, and the balance can still be overdrawn while every withdrawal clears.
Figures and dates in this piece are drawn from the public record of the Dutch childcare-benefits scandal and are dated where they move with the reporting. The two data-protection fines (2.75 million euros, decision 25 November 2021, announced 7 December 2021; 3.7 million euros, 7 April 2022) and their itemization are as reported by NL Times and the decision aggregator thedpo.eu against the regulator's own decisions; the recovery-cost tally (past 7.2 billion euros by 2024, up to a 14 billion projection) and the compensation figures are from NOS and De Correspondent reporting through 2024; the 2026 status is from Dutch News and Erasmus Magazine.
Sources:
The auditor in this story could see the whole model, say in public that it did not comply, and watch it keep running. The same gap opens wherever an agent decides at volume and the only record of what it did is the record it writes about itself. Chain of Consciousness keeps a provenance record of what an agent actually did, over traffic someone else can check, so the person who can see the aggregate is looking at evidence rather than at a summary.
pip install chain-of-consciousness npm install chain-of-consciousness
Or start without installing anything: Hosted Chain of Consciousness.