An APC teardown of the Stack Overflow Developer Survey. The numbers are real; the story about who was always optional, and one line of century-old arithmetic shows exactly why a survey like this can never pin it down.
In June 2022, Stack Overflow published its annual developer survey and handed the internet a headline it had printed six times before: Rust was the “most loved” programming language, for the seventh year in a row, with 87% of its users saying they wanted to keep using it. When the survey retired “loved” for the sterner “admired” in 2023, Rust just kept winning: 83% admired in 2024 (“for the second year in a row,” as Stack Overflow's own write-up put it), 72% and still #1 in 2025.
You have read the think-piece this figure generates. You may have written it. It goes: young developers love Rust. Juniors chase memory safety the way their elders chased garbage collection; the kids grew up on the borrow checker; Gen Z devs are built different. It comes with a chart, the chart goes up and to the right, and the conclusion feels like it is sitting right there in the data.
Here is the uncomfortable thing I want to show you, with arithmetic: that conclusion is not in the data. Not because the sample is too small, or self-selected (though it is), or because correlation is not causation. Something sharper. The dataset, any dataset shaped like this one, is mathematically incapable of telling you whether young people drive a trend. The proof is one line long, it is about a century old in demography, and once you see it you will spot undeclared versions of it in half the trend pieces you read.
The Stack Overflow survey is what statisticians call a repeated cross-section: every year, a fresh pile of respondents, each row carrying an age bucket and a survey year. From those two fields you can derive a third: roughly, birth year, or, closer to what tech punditry actually means, the year this person entered the field (the survey's years-coding fields give you that flavor too). Demographers call these three clocks age, period, and cohort, and every generational claim is a claim about which clock is doing the work:
The trap is that the three clocks are not three measurements. Cohort = period − age. Exactly. Know any two and you know the third, which means a model trying to estimate all three is asking the data to split one number three ways.
This is the age-period-cohort identification problem, and it is not a folk worry; it is a theorem about the geometry of the question. Put a variable for age, a variable for survey year, and a variable for cohort into one regression and the design matrix is rank-deficient by exactly one: its columns are linearly dependent, and infinitely many different coefficient vectors reproduce the observed data identically, not approximately but identically. More respondents do not help. A bigger survey sharpens every one of the competing answers equally and leaves them exactly as tied. The ambiguity lives in the question, not the noise.
Claims like that deserve a demonstration, so let us run one on the survey's cleanest curve. Stack Overflow started asking about AI tools in 2023, and the topline is famous: 70% of respondents using or planning to use AI tools in 2023, 76% in 2024, 84% in 2025. (The survey does not publish the age-by-year cross-tab for this at page level, so lay those real yearly margins across a small grid of age groups. The algebra we are about to watch does not care how the cells are filled, which is rather the point.)
Fit the linear age-period-cohort model and ask the standard software for the answer. Here is what comes back, three different runs, three different constraint choices, the kind of thing a careful analyst might try:
| age slope | period slope | cohort slope | worst fit error vs. the others | |
|---|---|---|---|---|
| Pseudoinverse solution | +2.33 | +4.67 | +2.33 | 0.000000000000 |
| “No period trend” constraint | +7.0 | 0 | +7.0 | 0.000000000000 |
| “No age, no cohort trend” | 0 | +7.0 | 0 | 0.000000000000 |
Read the rows as headlines. Row two says the trend is all age and cohort: the generational story, juniors and the class-of-2021, +7 points a step each. Row three says it is all the moment: ChatGPT year, everyone at once, no generation involved. Row one, the diplomatic compromise a fancy estimator produces, splits it 2.33 / 4.67 / 2.33 and looks, to the untrained eye, like a finding.
The last column is the punchline. Across every cell of the grid, the three fits differ by zero. Not “within confidence intervals.” Zero, to machine precision. They are the same surface wearing three stories, because the reallocation between them (add δ to the age slope, subtract δ from the period slope, add δ to the cohort slope) cancels exactly, for any δ, forever. That is the rank deficiency made flesh. I checked the matrix: four columns, rank three.
And that diplomatic-looking first row deserves its own paragraph, because it has a name and a fan base. In 2004, Yang, Fu and Land published the Intrinsic Estimator in Sociological Methodology, a principled-sounding fix that uses a pseudoinverse to pick, out of the infinite family of equally-fitting answers, the unique one orthogonal to the null space of the design matrix. It has lovely statistical properties. It is also, and this is the crucial part, a choice: one member of the tied family, selected by a criterion that has nothing to do with how developers actually adopt tools. In 2013, Andrew Bell and Kelvyn Jones published a paper in Social Science & Medicine whose title states the thesis with admirable bluntness: “The impossibility of separating age, period and cohort effects.” Their argument, backed by simulation: no estimator solves this, because the problem is inherent to the real-world process, not the statistics; the sophisticated methods recover the truth only when their hidden assumptions happen to match it. The fanciest estimator is not the one that found the answer. It is the one that best disguised the assumption.
The blog post that eyeballs a chart and says “kids these days” and the paper that runs an Intrinsic Estimator are making the same move. The blog is just easier to catch.
Here is where this gets genuinely useful rather than merely nihilistic, because the impossibility has a precise boundary, and the interesting action is right at the line.
Within one survey year, age comparisons are fine. If 2025's 18-to-24-year-olds admire Rust more than 2025's 45-year-olds, that is an observable fact, a cross-sectional gradient, no identification problem at all. What you cannot do is attribute the drift across years to any one clock.
Curvature survives. Straight lines don't. The unidentifiable piece is exactly the shared linear trend, the smooth drift. Kinks, spikes, accelerations, the second-difference structure, are identified, because reallocating a straight line among three clocks cannot manufacture or absorb a corner. In our AI curve, 70 to 76 to 84 is +6 then +8. The acceleration, +2 points, came out identical in every fit I ran, under every constraint. The near-vertical jump the year ChatGPT broke is real, attributable, and visibly a period shock: everyone, every age bucket, at once.
Sit with the irony of that. The one thing this dataset can actually pin down, “a moment moved everybody,” is the least generational story available. The moment a narrative becomes interesting (“this cohort is different,” “each class more than the last”) is precisely the moment it slides into the unidentifiable linear subspace. The seductiveness and the unprovability are the same mathematical property. A smooth generational drift is, by construction, the thing this survey can never attribute; a boring simultaneous lurch is the thing it can.
And there is a fourth clock nobody models. The survey's own respondent pool is sliding under the analysis: Stack Overflow's 2024 write-up notes that respondents aged 35 and up were 31% of the sample in 2022, 35% in 2023, and 39% in 2024. This is a self-selected convenience sample that re-draws itself every year. A composition shift like that can manufacture, mask, or reverse any apparent trend on any of the three clocks, before we even reach the theorem. The pundit's model does not just pick an unfalsifiable member of a tied family; it does so on top of a sample whose membership is quietly aging eight points in two years.
If this all sounds like an academic gotcha, watch what the people whose whole job is survey inference did when they finally stared at it.
In May 2023, Pew Research Center, the closest thing survey research has to a household name, published a methodological statement titled “How Pew Research Center Will Report on Generations Moving Forward.” In it, they concede the core point: when younger adults answer differently than older ones, “it may be driven by their demographic traits rather than the fact that they belong to a particular generation.” Their new house rules: generational claims require decades of comparable historical data, and where the label is not earned, they will group by decade or event instead. They also described the generational-content industry, with visible fatigue, as a crowded arena where much of what is “sold as research” is closer to marketing mythology.
A century of theory sits behind that retreat. Karl Mannheim's 1928 essay “The Problem of Generations,” still the founding document of the field, was already more careful than the genre it spawned: he argued that merely sharing birth years creates only a potential for shared consciousness, that actual “generation units” form around formative experiences in the impressionable years of late adolescence, and that no generation is a homogeneous block. The modern methodological literature (Bell and Jones among them) supplies the theorem underneath his caution: the clean version of the claim everyone wants to make is not just hard. Without an outside assumption, it is unavailable.
So when a survey house of Pew's stature stops writing “Gen Z believes…” headlines from single-year cross-sections, that is not timidity. That is a detector being honest about what it can detect, while much of tech commentary keeps confidently reading the one component of the signal that is provably unreadable.
The practical insight is not “never trust surveys.” It is a one-question audit you can run on any trend claim, others' or your own, in about ten seconds:
“Which clock did you zero, and where did you say so?”
Every attribution of a repeated-cross-section trend to age, generation, or moment has zeroed at least one clock. There is no exception; the algebra does not permit one. The only variables are whether the author knows they did it, and whether they told you. From there, three habits:
And if you are the one writing the analysis: say the null proudly. “This survey cannot tell whether juniors drive Rust adoption” is not a failure to find a result. It is the result: a falsifiable claim that happens to invalidate a whole genre, and the strongest sentence the data will underwrite. The numbers are real: 87%, seven years running, 83%, 72%, 70-76-84. What is optional, what was always optional, is the story about who. Rust may well be beloved by the young. The Stack Overflow survey, read honestly, can neither confirm that nor deny it; it can only watch the whole field move and decline to say which clock did it.
Anyone who tells you otherwise has made an assumption. The good ones tell you which.
survey_results_public.csv, survey_results_schema.csv) published under the Open Database License.Every attribution of a repeated-cross-section trend has zeroed at least one clock. The only question is whether the author told you which.
That is the discipline Chain of Consciousness brings to an agent's decisions: a tamper-evident record of what the agent actually considered and why, so the assumption behind a conclusion travels with the conclusion instead of living in someone's head. When the reasoning is part of the artifact, “which clock did you zero, and where did you say so” stops being a question you have to trust an answer to and becomes one you can check. It is one layer of the Agent Trust Stack, the harness for making agent behavior verifiable, claimable, and auditable rather than taken on faith.
See Hosted Chain of Consciousness · Read the Theory of Agent Trust
pip install chain-of-consciousness · npm install chain-of-consciousness
Full trust stack: pip install agent-trust-stack · npm install agent-trust-stack