Does the School or the Country Teach Your Kids? Two Decades of PISA, Recomputed

Published September 2026 · 11 min read · PISA / education / variance decomposition / statistics

In the late 2000s, education ministers flew to Helsinki the way engineers now fly to conferences about AI. Finland had posted a 548 in PISA mathematics in 2006, near the top of the world, with almost no standardized testing, short school days, and no school-inspection regime. Delegations toured its classrooms looking for the secret. In December 2023, the Finnish public broadcaster Yle ran the other headline: "Performance in Finland collapses." The 2022 assessment put Finnish 15-year-olds at 484 in math, 64 points below the 2006 peak, 23 of those points lost since 2018 alone.

So Finland fell. But fell at what, exactly? A country's PISA score is a single number stretched over three very different things: the country you're born in, the school you attend within it, and the household you go home to. When the headline number moves, it almost never says which layer moved. And the layers are not the same size, they do not move at the same speed, and only some of them are things a parent, a teacher, or a ministry can actually act on.

The OECD publishes the microdata for every cycle back to 2000: hundreds of thousands of students per wave, each with a math score and a socioeconomic index. The full decomposition of that data is a real statistical project, and this piece is honest about not having run all of it. But a useful slice of the question can be computed directly from public data in an afternoon, and the published record fills in the rest. I did the afternoon's worth. Here is what the layers look like when you take them one at a time.

Layer one: the league table barely moves

Start with the layer that gets all the coverage, the country rankings. The World Bank republishes the PISA mathematics country means as an open data series, and its mirror covers the seven cycles from 2000 through 2018. I pulled that series through the public API and computed the thing the headlines imply must be turbulent: how much the league table actually reshuffles from one cycle to the next.

It barely does. Taking the countries present in each adjacent pair of cycles and correlating their rank orders, the Spearman correlation never drops below 0.93, and in most windows it sits between 0.965 and 0.987. The average country moves between 1.5 and 3.3 rank positions per cycle. The top ten loses at most two members per cycle, and in the 2006-to-2009 window it lost none at all. The exits that did happen are a short and specific list: the United Kingdom after 2000, Belgium after 2003, Canada and Finland after 2009, Finland again from the 2012 list, Canada after 2015.

Even across the long span, stability dominates: the 2003 and 2018 rank orders correlate at 0.865 across the 39 countries present in both. The league table that gets reported every three years as an earthquake is, statistically, one of the most stable rankings in public data.

What does change, slowly and legibly, is a handful of long drifts. Over 2003 to 2018, the biggest declines in the series are Finland (down 37 points), Australia (down 33), New Zealand (down 29), and Belgium (down 21); the biggest gains are Macao and Türkiye (up 30 each), Brazil (up 28), and Portugal (up 26). Compound those drifts over two decades and the top ten's composition transforms: in 2000 it held five Anglosphere systems plus Finland; by 2018 it was East Asian systems plus Estonia, the Netherlands, Poland, and Switzerland. The reshuffling is real, but it is a twenty-year tide, not a three-year churn.

Two honesty notes on this computation. The country means are point estimates, and neighboring countries often sit within each other's measurement error, which is exactly why I computed rank correlations, a statistic that is coarse enough to survive that noise, rather than celebrating any single country's two-place move. And the World Bank mirror stops at 2018 for this indicator, so the 2022 cycle enters this piece only through the OECD's own publications.

Layer two: the school matters less than you think

Now the layer parents lose sleep over: which school. PISA's own analysts decompose score variation into a between-school component and a within-school component, and the OECD's report on the 2022 cycle states the split plainly: across OECD countries, 32% of the variation in mathematics performance lies between schools, and 68% lies within them.

Read that again from a parent's chair. If you could magically swap your child among schools while holding everything else constant, you would be operating on the one-third of the variation, at most, that school membership can carry. Two-thirds of the spread in outcomes exists inside each school building, between the kids who walk the same corridors.

And the between-school share is not a law of nature; it is a policy choice you can see in the numbers. Systems that sort children into academic tracks early show large between-school variance, because the sorting manufactures it. Comprehensive systems compress it. In the 2009 cycle, less than 10% of Finland's variance sat between schools, against 30-35% in several East Asian systems, a contrast documented in the World Bank's analysis of that wave. In Finland, which school you attended told you almost nothing about how you'd score. That was the actual Finnish miracle, and it is a different miracle than the one the delegations went looking for.

Layer three: the household, the stickiest layer of all

PISA attaches to every student an index of economic, social, and cultural status, built from parental occupation, parental education, and home possessions. Regress scores on that index and you get the socioeconomic gradient, the statistical shadow of the family behind the student.

The published record on this layer is remarkably consistent. In the 2015 cycle, the index explained about 13% of score variation on average across OECD countries, ranging from around 5% in Iceland and Hong Kong to more than 20% in France and Peru. In 2022, the OECD reported the gap between the top and bottom socioeconomic quartiles at 93 points of mathematics on average, with disadvantaged students seven times more likely to miss baseline proficiency. A 2024 systematic review in the International Journal of Educational Research Open surveyed twenty years of studies on exactly this question and concluded that, by the OECD's own metrics, the goal of more equitable systems remains largely unfulfilled: the gradient varies across countries, but the overall picture over two decades is stubborn stability.

Set the three layers side by side and the original question answers itself, with a twist. Does the school or the country teach your kids? Mostly neither. The country means move slowly and mostly in a few famous cases; the school explains a minority of the variation, and less than that in systems designed to compress it; and the layer that has resisted twenty years of reform effort in nearly every system is the one no ministry controls, the household. The headline layer is the most volatile and least decisive; the household layer is the most decisive and least movable. The public conversation runs in exactly the wrong order.

The Finland test

Here is where the three layers earn their keep, because Finland's fall looks completely different depending on which layer you read it in.

In layer one, Finland is a catastrophe: minus 64 points from peak, out of the top ten, headlines about collapse. In layer two and three, something surprising: in the OECD's 2022 equity analysis, Finland still appears on the short list of ten systems that combine broad baseline proficiency with high socioeconomic fairness, alongside Canada, Denmark, Japan, Korea, Ireland, Latvia, the UK, Hong Kong, and Macao. The country whose mean collapsed is still one of the places where which school you attend, and which family you come from, predicts least about how you'll do relative to your peers.

That is not a consolation prize; it is a diagnosis. Whatever pulled the Finnish mean down over sixteen years, and the national post-mortems point at everything from phones to demographics to curriculum drift, it pulled the whole distribution down together. The structure that made Finland famous, the flat, comprehensive, low-sorting system, is still there and still doing what it did. A mean is not a mechanism. If you read only the league table, you would prescribe Finland a revolution in exactly the layer that isn't broken.

The same test runs in reverse on the systems that climbed. When a country rises thirty points over fifteen years, the league table cannot tell you whether it lifted its median student, stretched its top track further from its bottom track, or changed who sits the test. The decomposition can. That is the whole argument for computing it per cycle instead of screenshotting the rankings.

How you'd compute the real thing, and why it's harder than it looks

Everything above uses one public series plus the OECD's published decompositions. The full version of this investigation, the per-cycle, three-level variance decomposition with the gradient tracked across twenty years, is sitting in the open, and the route to it is short. The OECD's file server hands out the microdata with no login: for the 2022 cycle, the student questionnaire file carries the socioeconomic index and the school identifier, the cognitive file carries the scores, and a school file carries school characteristics. Three downloads and you hold the raw material for every claim in this essay.

Then four traps close around anyone who computes carelessly, and each one has a lesson that generalizes far beyond education data.

First, PISA does not give each student a score. It gives each student ten plausible values, draws from a probability distribution over what that student's proficiency might be, because a teenager who sat one test booklet is a noisy measurement. Compute your statistic on each of the ten and combine them with Rubin's rules, or your uncertainty estimates are fiction. Anyone who has shipped an eval that reports a single accuracy number from a single sampled run has met this trap wearing different clothes.

Second, the survey's standard errors come from balanced repeated replication with dozens of replicate weights, not from textbook formulas. Skip them and your error bars shrink to a fraction of their honest size, which matters most for the central claim here, because "the gradient did not change" is a null claim, and a null claim with fake-narrow error bars is unfalsifiable theater.

Third, the socioeconomic index gets rebuilt and rescaled between cycles. A gradient slope in 2000 units is not a slope in 2022 units until you've handled the comparability explicitly. Metrics drift under stable names everywhere; the index that kept its name did not keep its meaning.

Fourth, the set of participating countries changes every cycle. A twenty-year trend computed over whichever countries happened to show up is comparing different worlds. You hold the panel fixed, and say so. My rank correlations above dodge this by construction, pairwise, but a pooled variance trend cannot.

None of these traps is exotic. They are the same four failures that quietly invalidate dashboards: point estimates passed off as certainties, uncertainty computed wrong, metrics redefined mid-series, and populations that shift under the trend line. PISA is just a rare case where the survey's designers documented all four and handed you the machinery to do it right.

What to do with this

If you're reading as a parent: the school-choice decision carries less predictive weight than the anxiety around it suggests, one-third of the variation at the outside, and in comprehensive systems far less. The layer with the largest stable association is the one you already run: the household. That is either comforting or damning, but at least it is quantified.

If you're reading as someone who ships metrics for a living: the PISA league table is a twenty-year masterclass in how a headline number gets read in the wrong layer. Before anyone reorganizes around a moved metric, decompose it. Ask what share of the movement is between groups, what share is within them, and whether the gradient that actually determines individual outcomes moved at all. The league table reshuffled slowly; the structure underneath it barely moved; and the loudest collapse of the era happened in a country whose equity machinery, on the OECD's own 2022 numbers, is still among the best on earth.

The full recomputation is one afternoon of downloads and one careful week of statistics away, and the four traps above are the map. Until someone runs it per cycle and publishes the decomposition, every three years the same ritual will repeat: the rankings will land, a country will be declared broken or brilliant, and the number doing the declaring will be the one layer of the three that says least about any actual child.


Sources: OECD, PISA 2022 Results, Volume I (2023), performance and equity chapters, including the 32%/68% between/within-school split, the 93-point quartile gap, and the ten-system fairness list; OECD/Finnish Ministry of Education country materials for Finland's 484 in 2022 and the 64-point decline from 2006; Yle News, December 2023; World Bank Open Data, indicator LO.PISA.MAT (PISA mathematics means, 2000-2018), from which the rank-correlation, rank-movement, and top-ten-churn figures above were computed by the author this week; World Bank / NCBI, "Global Variation in Education Outcomes at Ages 5 to 19" for the 2009-cycle between-school shares; OECD, PISA 2015 Results for the 13% ESCS variance-explained average; "Change in socioeconomic educational equity after 20 years of PISA: a systematic literature review," International Journal of Educational Research Open (2024); OECD PISA microdata file server (webfs.oecd.org) for the data route described.


If you publish a score that other people act on, the layer problem is yours too. A single rating compresses provenance, sample, and scope into one number, and the people reorganizing around it cannot see which of those moved. Agent Rating Protocol keeps a rating attached to what it was computed over, so a score can be decomposed by whoever reads it rather than taken at face value.

pip install agent-rating-protocol
npm install agent-rating-protocol

See how a decomposable rating is verified