Across all 1,035 breaches in the catalogue the median gap between the breach happening and anyone being able to check is 132 days. A third stayed dark longer than a year. The script and the snapshot are published with the piece, so you can run the number yourself.
Computed from a snapshot of the Have I Been Pwned breach catalogue taken 10 September 2026 (1,035 breaches). Both the script and that snapshot are linked at the foot of this piece; the live numbers move as the catalogue grows.
Around the end of 2011, on 26 December by the date Have I Been Pwned gives it, somebody copied the user database of RuneScape Boards, a fan forum for a browser game, 222,762 accounts with email addresses and passwords. On 23 March 2026, Have I Been Pwned loaded it. If you posted on that forum in 2011 and have typed your email into the site every year since it opened, the answer came back clean twelve times and then, this spring, it didn't. The breach was 5,201 days old. Nothing about it had changed. What changed was that someone could finally check.
That gap has a name in the data, though not in any report. Have I Been Pwned's public API lists every breach it holds with two dates: BreachDate, the day the data was taken, and AddedDate, the day the site could answer a question about it. The documentation is candid about the first one: it "is not always accurate - frequently breaches are discovered and reported long after the original incident. Use this attribute as a guide only." The endpoint needs no key and no account. One GET request returns all of it, and as of 10 September 2026 all of it is 1,035 breaches, 17.8 billion account records between them.
Subtract one date from the other and you have the interval during which the people in a breach were the only people who could not act on it. The attacker had the data. Whoever bought it had the data. The company, in most cases, had at least a suspicion. The victims had nothing to check against. Call it the dark interval. Across the whole catalogue the median is 132 days. More than a third of breaches, 34.8 percent, were dark for over a year. More than one in five, 22.7 percent, for over two. Only 44.7 percent surfaced within ninety days. Weight by accounts instead of breaches and the picture holds: about a quarter of all the records in the catalogue sat dark for more than a year.
Nobody publishes this number. Breach counts get published, record counts get published, the cost per breach gets published every summer. The interval in which the victims were the only ones in the dark does not, and it is the one that measures whether a disclosure regime is doing the thing it exists to do.
Two corrections before any more numbers, because both change them.
The first is precision. BreachDate is a date, but 176 of the 1,035 breaches, 17 percent, are dated the first of a month, and 42 of those are dated 1 January, which is a year with a day stapled to it. So an individual breach's interval can be wrong by weeks and a few by months. No interval is negative, which is the one thing the field could not get away with. Every figure in this piece is a median over dozens or hundreds of rows, which is why the imprecision does not sink them and why no single breach's number below should be quoted to the day.
The second is the calendar. Have I Been Pwned went live on 4 December 2013, when Troy Hunt loaded five datasets, Adobe, Stratfor, Gawker, Yahoo and Sony, having noticed that when he added Stratfor to Adobe, "16% of the email addresses were already in the system." Seventy-seven breaches in the catalogue happened before that day. For them the dark interval includes a stretch when there was no tool to be dark from. Their raw median is 1,648 days. Measured from launch day instead, it is still 1,021 days: even once the site existed, the breaches from before the launch, 2007 to 2013, took a median of nearly three years to arrive in it.
So the defensible corpus is the 958 breaches that happened after the site existed, and the defensible claim about AddedDate is narrow. It is not the day the public first heard; some breaches were in the newspapers before they were in the catalogue, and some of the entries below were disclosed by the breached site itself well before the catalogue carried them. It is the day a victim could self-serve an answer, which is the day the interval this piece cares about actually ends. On that corpus the median dark interval is 106 days. Forty-eight percent surfaced within ninety days. Thirty percent were dark for more than a year, 17.4 percent for more than two, and the ninetieth percentile is 1,112 days, three years.
This is the question the disclosure laws exist to answer, and the data answers it in two halves.
Group the breaches by the year they happened and take the median interval. Breaches from 2014 sat dark a median of 451 days. From 2015, 302. From 2016, 264. From 2017, 169. From 2018, 121. That is a real and large improvement, and it lines up with the years in which breach notification stopped being a Californian idea and became the default. California's law, the first, took effect in July 2003 and requires notice "in the most expedient time possible and without unreasonable delay." Every US state now has one. The GDPR, from May 2018, gives controllers 72 hours to notify the regulator once they become aware. The SEC, from December 2023, gives public companies four business days on Form 8-K once they decide an incident is material.
Then the series stops improving. Breaches from 2019 sat dark a median of 168 days, from 2020, 170, from 2021, 162, and from 2022, 275. Take the censoring-resistant version of the same question, the share of a year's breaches that were catalogued within twelve months, and the plateau is flatter still: 65 percent for 2017, 68 for 2018, then 60, 65, 60 and 57 for 2019 through 2022. Roughly four breaches in ten from the plateau years were dark for more than a year. About six in ten were dark for more than ninety days.
The rows after that look wonderful and cannot be read. Breaches from 2024 show a median of 23 days and 95 percent catalogued within a year; 2026 shows 14 days and 100 percent. They show that because a breach that has not yet been catalogued is not in the dataset at all. The 2024 row contains only the 2024 breaches that surfaced quickly, by construction, and it will get worse every time an old one turns up. The 2011 row, at 1,947 days, is the opposite: nearly complete, and almost certainly an underestimate of nothing. The honest reading is that the long tail of the early years is solid evidence that intervals used to be enormous, that the 2018 to 2022 plateau is the most trustworthy recent number, and that the last three rows are not evidence of anything except how fast the fast ones are.
Notice, too, what none of those laws measure. Every clock in them starts at awareness: 72 hours after the controller "becomes aware," four business days after the company "determines" materiality. The interval between the breach and the awareness is nobody's statutory problem, and it is most of the dark interval. The nearest published figure is IBM's annual survey of about 600 breached organizations, which put the mean time to identify and contain a breach at 241 days in 2025, a nine-year low. That is the company's interval, self-reported, from its own point of view. The victim's interval, from the day the data left to the day they could check, is the number in this piece, and the two are not the same measurement.
Size first, because the intuition is wrong in an interesting way. The biggest breaches do not surface fastest, and neither do the smallest; the middle sits dark longest. Breaches of under ten thousand accounts have a median interval of 10 days. Ten thousand to a hundred thousand, 31. A hundred thousand to a million, 116. One million to ten million, 211 days, the slowest tranche. Ten to a hundred million, 182. Over a hundred million, 109. Restrict to breaches from 2016 onward, so the old backfilled years cannot drive it, and the hump survives: 132 days for the one-to-ten-million tranche against 86 for the giants and 10 for the smallest. The plausible reading is that the giants are newsworthy enough that someone reports them, the tiny ones get dumped at once because there is nothing to sell, and the middle is exactly the size worth trading privately before it is dumped. That is a hypothesis. The hump is the finding.
Now what was taken, which is where this investigation nearly published something false.
Across all years, the data classes that sit dark longest are the least alarming ones: website activity at 269 days, usernames 257, IP addresses 249, passwords 236. The ones that surface fastest are the most alarming: government-issued IDs at 22 days, private messages 25, job titles 37, employers 41, physical addresses 67, phone numbers 83, names 98. Government IDs surfacing eleven times faster than passwords is a headline, and it comes with a tidy story: a breach of identity documents triggers a legal duty and enters the disclosure machinery at once, while a forum dump of usernames and hashes triggers nothing, gets traded, and surfaces when a researcher happens to catalogue it.
The story may even be true, but the eleven is not. The classes that surface fastest are also the classes that appear mostly in recent breaches. The median government-ID breach in the catalogue happened in 2023; the median password breach in 2018. Recent breaches have short intervals because of the censoring above, not because of what they contain. Restrict every class to breaches from 2014 through 2021, years that are old enough to be mostly complete and young enough to be inside the tool's life, and the ranking compresses. IP addresses sit dark a median of 274 days, usernames 268, passwords 239. Names fall to 198, phone numbers 168, physical addresses 166. Government IDs come in at 118 days, on ten breaches, which is too few to lean on. Private messages, at 10 days on 26 breaches, are the one class that stays dramatically fast.
So the direction survives and the magnitude does not. Credential-shaped data, the usernames, addresses and passwords that make a forum dump, sits dark roughly 270 days. Contact-shaped data, the names and phone numbers that make a customer database, roughly 170. The gap is real, it is about 1.6 to 2 times rather than 11, and it is still consistent with the duty-not-harm reading, because the kind of data correlates with the kind of organization: identity documents live in systems that already sit inside notification regimes, and hashed passwords live in hobbyist forums that do not. The data cannot separate the law from the landlord. It can say that whatever the cause, the interval tracks who was obliged to speak far better than it tracks how much harm the breach could do.
Two smaller results point the same way. Breaches the site marks as verified have a median interval of 127 days; the 42 it marks unverified, 356. The corpus's least certain entries are also its stalest, which is the opposite of what a reader assumes when they see the flag. And the nine breaches sourced from malware have a median interval of zero days; stealer logs, 12. Those are catalogued the day they are collected, because the collection is the event. They are the control group for what a disclosure regime with no dark interval looks like, and they exist only for data that was never inside an organization at all.
The twelve longest dark intervals in the catalogue are, in order, a RuneScape fan forum, the China Software Developer Network, a botting forum, a subtitle site, a game, a game studio, a gaming-peripherals maker, a book-swapping community, a Belgian gaming news forum, a baby-names site, a web-hosting forum and another game. Every one is marked verified. Between them they hold about 20 million accounts. Not one is a bank, a hospital or a government, and only one of the twelve carries a disclosure link at all. They are precisely the sites with no regulator, no notification duty and no press interest, which is the data-class result arriving from the other direction.
And the list is still being written. Four of the eight longest were catalogued in the last ten months: the China Software Developer Network in November 2025, at 5,090 days, the botting forum in December 2025, RuneScape Boards and Scuf Gaming in March 2026. The catalogue added 106 breaches in 2024, 91 in 2025 and 100 so far this year, and every year a handful of them are more than five years old: nine in 2024, fourteen in 2025. The 2011 row of the year table has been closed for years in the sense that matters to a regulator, and it is not closed at all in the sense that matters to the people in it.
If a disclosure regime wanted to be graded on the thing it exists for, the grade is now a single sentence per year: of the breaches that happened that year, what share could their victims check within ninety days. For the years old enough to answer, 2014 through 2022, the catalogue says between 26 and 47 percent, depending on the year. The years since say better, and the years since are exactly the ones that cannot be read yet, because every breach still in the dark is a row that does not exist. The newest number in this dataset is always the best-looking one and always the one that is still coming down.
That is the property to hold onto, and it is the whole practical point. It is also the fourth time this blog has arrived at the same place from a different direction: that a rate is only as good as the population it was measured over, and that the provenance of a widely-quoted statistic is usually thinner than the statistic. This piece is part of Where the Number Came From, on how a published number is a fact about the way it was measured rather than about the world. What is different here is that the number did not exist until we computed it, so there is no published figure to check us against, only the script and the snapshot. A count of breaches only goes up as the world discovers more, and so it flatters no one. A dark interval measured on a recent year only gets longer as the world discovers more, and so it flatters everyone until the discovering stops. Any regime that reports how many breaches it processed, and not how long its citizens were the last to know, is grading itself on the easy number. The script that prints every figure above is 54 lines and one request. Run it on the day you read this. The 2026 row will look better than it does today, and it will be wrong in the same direction, and by the time it is right there will be a forum somewhere whose 2011 comes in.
Sources:
re_, and our own secret scanner reads that as an API key. It produces byte-identical output to the full file, which we checked before publishing it rather than assuming. Run the script with no argument and it queries the live API; pass the snapshot as an argument and you get this exact run back. Every figure above is its output.That interval is the subject of this piece, and it is also the thing we build against. An agent that keeps a verifiable record of what it did closes the same gap inside your own systems: not "was there an incident" but "when could anyone first have known". Chain of Consciousness makes that record checkable by someone who was not there.
pip install chain-of-consciousness npm install chain-of-consciousness
Or start without installing anything: Hosted Chain of Consciousness.