← Back to blog

Six Users on Day One: A Teardown of the Healthcare.gov Rescue, Not the Failure

The famous number was right. So was the 1,100-user bulletin, and the volume diagnosis. Every layer of this story read a real number off the wrong instrument.

Published August 2026 · 10 min read · observability / measurement / incident review / government software


On the morning of October 2, 2013, in a meeting room at the Centers for Medicare and Medicaid Services, somebody wrote down a number. The note, from the war-room minutes of the office running the brand-new HealthCare.gov, reads: "6 enrollments have occurred so far with 5 different issuers." Those minutes, released a month later by the House Oversight Committee, became the most famous number in the history of government software. A five hundred million dollar website, and six people got through.

Here is what almost nobody quotes: the words "so far." That was the morning reading. By the afternoon meeting the same notes record roughly one hundred enrollments. By the next morning, 248. Around forty thousand applications were pending by the end of day two. No official day-one figure was ever produced, a point the Oversight release itself concedes, and PolitiFact, rating the "six people" claim Mostly True, carried the same caveat: initial numbers, no official count, paper and state enrollments not included.

So the six is not false. It is something more interesting than false. It is the first reading of a rapidly climbing counter, frozen at its lowest value and made permanent. And the instrument that produced it was not a dashboard or a metrics pipeline, because the site had neither. It was a meeting. People in a room, saying numbers out loud, writing them by hand.

Hold onto that image, because this whole story, the disaster and the rescue and the investigation of both, is a story about instruments. Everyone in it was reading one. Most of the numbers were right. The scandal, at every layer, is what they were measuring.

A system nobody could see

Start with what the site actually shipped without. The U.S. Digital Service, in its later report to Congress, states it in one flat sentence: "At initial launch of HealthCare.gov, there was no end-to-end monitoring of the production system, making identification, prioritization and diagnosis of errors very challenging."

CMS's own December 2013 progress report says the same thing from the inside: "The system monitoring and response mechanisms were not sufficient for identifying issues or bugs or responding to them in real time." And it describes the user experience in a phrase that deserves more attention than it gets: consumers were seeing "frequent, inexplicable error messages."

Inexplicable. In the government's own report. Not inexplicable to users, who never expect an explanation. Inexplicable to the operators. The people running the system could not explain its errors, because the thing that explains errors, instrumentation, had not been built. During some weeks of October, by CMS's own accounting, the site was down an estimated sixty percent of the time, and the team inside had roughly the same view of why as the public did: none.

That is failure layer one, and it is the ordinary kind. Teams ship without observability constantly. The dashboard gets written during the first outage, by the light of the fire. What makes HealthCare.gov worth a teardown is what happened next, twice.

The diagnosis that outran its instrument

Within days of launch, senior officials explained the failure to the country. On October 2, the White House press secretary: "There's no question that the volume was so high and continues to be so high that that has caused some delays. Those delays are, in our view, related to the high volume." On October 6, the U.S. Chief Technology Officer, Todd Park, was more specific: officials had expected around sixty thousand simultaneous users and drew roughly two hundred fifty thousand. "These bugs were functions of volume," he said. "Take away the volume and it works."

Read that sentence next to the USDS finding. "Take away the volume and it works" is a counterfactual causal claim about a system that, at that moment, had no end-to-end production monitoring. It is not a lie. It is worse in a subtler way: it is a story, told confidently, by people whose instruments could not have licensed it. And six weeks later, CMS's own report quietly contradicted it, attributing the failure to "hundreds of software bugs, insufficient hardware and infrastructure." Volume was the trigger. The list of causes was long, and the site's inability to see itself sat near the top.

Every engineering team knows this move, because every engineering team makes it. The incident starts, the graphs are useless or absent, and within the hour somebody with seniority says it's the database, or it's a thundering herd, or it's the deploy. The tell is always the same: a causal claim whose only evidence is a coincidence in time. Layer two of this story is just that move, performed at podium height.

The investigators make the same error while investigating it

Here is the part of the story I have never seen told, and it is the reason this essay is about the rescue's epistemics rather than the launch's embarrassments.

In November 2013, the House Oversight Committee, investigating exactly the failures above, released a contractor's daily testing bulletin under a national headline: HealthCare.gov could only handle 1,100 users the day before launch. The bulletin is real, dated September 30, written by QSSI, and it says what the committee said it says: "Currently we are able to reach 1100 users before response time gets too high."

It also says, in the same passage, that the test ran "in IMP1B environment."

IMP1B was a test environment. Not production. The next day, the committee's ranking member released portions of a transcribed interview with Henry Chao, the CMS deputy CIO, whose account was that testing projected the full production system would handle up to fifty-eight thousand concurrent users, and that the chairman had, in the minority's words, "conflated the results of a much smaller testing environment with final production testing of the system at full capacity."

Both releases are partisan documents, and this essay is not going to referee them. It does not need to. Sit back and look at the shape. The bulletin's number was genuine, correctly measured, honestly reported by the contractor. It was then carried into a claim about a different system, and it ran nationally. Which means the committee investigating a site that shipped without working instruments, and officials who diagnosed it beyond what their instruments could support, produced its own headline by reading a real number off the wrong instrument.

Three layers. The operators could not see the system. The executives explained it anyway. The investigators audited both with a measurement from an environment that was not the one in dispute. Nobody in this story fabricated anything. That is precisely what makes it useful: every failure here is available to honest people, and most of them are available to you.

What the rescue actually did first

The rescue is remembered as a heroic engineering sprint, and the folklore version stars brilliant outsiders rewriting a broken site. What the primary record shows is quieter and more instructive: almost nothing got rewritten, and the first thing the rescue team did was not code at all.

From the USDS report: "One of the first actions the 'tech surge' team took was to recommend the addition of an application monitoring tool, which has remained an important resource for the team to identify issues as they occur." First action. Before fixes, sight.

CMS's December report describes the second move: a single command structure, with QSSI appointed as general contractor and systems integrator, and "'War Room' meetings of all key parties held twice a day for real-time, data-based decision making."

Notice the recursion, because it is almost too neat. The fix for a system nobody could see was a twice-daily meeting where people read numbers aloud. That is the same instrument that produced the infamous six, five weeks earlier. The meeting was never the problem. A meeting is a fine instrument. What changed is that the numbers being read aloud were finally worth reading, because a monitoring layer now existed underneath them.

And then, with sight restored and a decision structure attached to it, the same system that had been down sixty percent of the time in October reached the December 1 progress report handling fifty thousand concurrent users with an error rate below one percent. The workflow rebuild that followed took application conversion from roughly fifty-five percent to eighty-five percent, and by March 2014 more than eight million people had enrolled, 5.3 million of them through HealthCare.gov. The site the six-enrollment note described and the site that did that were, in most of the ways that matter, the same code. What had been added was the ability to see it and a structure that acted on what was seen.

One more instrument note from the wreckage, because it is the failure mode most teams actually have. The Government Accountability Office found that the operational readiness review, the institutional gauge that answers "is this safe to launch," slipped to September 2013, weeks before go-live. The instrument existed. It was scheduled too late for its reading to change anything. A gauge you consult after the decision is decoration.

The discipline, in five lines

For working engineers and the people who lead them, the whole record compresses to five rules, every one of them purchasable for less than a national embarrassment.

First: observability is a launch requirement, not a maturity milestone. If your first dashboard gets built during your first outage, you have chosen the HealthCare.gov sequence voluntarily.

Second: an explanation produced without an instrument is a story. Stories are selected for plausibility, and plausible is what wrong explanations are made of. Before you accept "it's the volume," ask what measurement, specifically, licenses the claim, and whether it exists.

Third: a correct number from the wrong environment is more dangerous than no number, because it arrives wearing the costume of evidence. Every staging benchmark quoted as production capacity, every canary read as the fleet, every load test that exercised a warm cache is the 1,100-user bulletin with your logo on it.

Fourth: a single reading of a moving quantity is not a measurement of the quantity. Six and 248 are the same system thirty hours apart. If a number matters, ask for its second reading before you let it become a headline, in either direction.

Fifth: when a system cannot be seen, the first fix is organizational as much as technical. The rescue's twice-daily war room with a single empowered integrator did as much as the monitoring tool it was built around. Sight plus a structure that acts on sight; neither works alone.

The postmortem that produced an institution

The durable output of the rescue was not the website. In August 2014, the administration founded the U.S. Digital Service, staffed in its first round from a thousand applicants for ten positions, an agency whose founding logic traces straight back to a government discovering it could not see its own software. Most postmortems produce a document. This one produced a permanent capacity for the next emergency, which is the rarest and best corrective action there is.

But the institution is the epilogue. The lesson is the note. Somewhere in your organization, right now, there is a number moving through meetings the way "6 enrollments have occurred so far" moved through that one: a snapshot of a moving counter, taken off the only instrument available, about to be frozen into a fact. The people who wrote that note did nothing wrong. They read the instrument they had. The discipline this story teaches is to ask, before the number hardens, the question nobody asked in October 2013, and the investigators did not ask in November, and the rescue finally answered in December.

Not "what is the number?" The number is usually right.

"What, exactly, was the instrument measuring?"


Sources: House Oversight Committee release quoting the CCIIO war-room meeting notes of October 2 and 3, 2013 (October 31, 2013); House Oversight Committee release quoting the QSSI ACA Daily Testing Bulletin of September 30, 2013 (November 7, 2013); Oversight Committee Democrats release quoting the transcribed interview of CMS Deputy CIO Henry Chao (November 8, 2013); CMS, "HealthCare.gov Progress and Performance Report" (December 1, 2013); GAO-14-694, "Healthcare.gov: Ineffective Planning and Oversight Practices Underscore the Need for Improved Contract Management" (July 2014); U.S. Digital Service, Report to Congress 2016, "Stabilizing and Improving HealthCare.gov"; U.S. Digital Service, Mission (founding, August 2014); PolitiFact, rating of the "six people signed up on day one" claim, Mostly True (November 3, 2013). Both House releases are partisan documents cited here for the contemporaneous records they quote verbatim, not as neutral narration; this essay does not adjudicate their dispute.

An explanation produced without an instrument is a story

The rescue's first action was not a fix. It was adding the ability to see what the system had actually done. AI agents have the same gap: the output arrives, the reasoning does not, and the explanation gets reconstructed afterwards by whoever is asked. Chain of Consciousness is a tamper-evident record of what an agent did, on what inputs, in what order, written as the work happens rather than assembled once somebody wants an answer. It does not make an agent right. It makes the basis of its answer checkable, which is the difference between a reading and a story.

Hosted Chain of Consciousness  ·  Verify a record

pip install chain-of-consciousness  ·  npm install chain-of-consciousness