A saliency map that survives randomising the model was never about the model. The Price equation cannot be false, and for exactly that reason cannot explain. Same trap, three fields, one test you can run this afternoon.
In 2018 Julius Adebayo and five colleagues took a trained image classifier and destroyed it slowly. They randomised its weights one layer at a time, from the output backward, until the network was noise, and after each step they asked the standard explanation tools to show them why the model saw a bird, a dog, a junco on a branch. Several of the tools kept producing the same picture. The heat map that explained the trained model's decision explained the ruined model's decision too, with the same bright outline around the same bird. Then the team trained a fresh model on shuffled labels, so that it could not have learned anything about birds, and asked again. Same picture.
The paper's interpretation is the sentence to keep: some widely used saliency methods behave like an edge detector, "a technique that requires neither training data nor model." The explanation was a property of the input image and the method, not of the thing it claimed to explain. It would have been produced whatever the facts were.
That is the shape of a specific epistemic trap, and it is old. The oldest and cleanest version of it lives in evolutionary biology, in a one-page paper from 1970, and the argument about it is still running. Anyone who has ever written "the model does X because feature Y" and been proud of how tidy that sentence was has a stake in it.
George Price's 1970 paper in Nature partitions the change in a population's average trait between generations. One term is the covariance between the trait and fitness, the amount of change that comes from individuals with more of the trait leaving more descendants. The other is transmission, the change that comes from the trait shifting on its way from parent to offspring. Selection, in Price's equation, is the covariance term. The equation holds for any population and any selection-like process, genes, cultures, firms, code, because it is not a claim about any of them. It is an identity: the right-hand side is the left-hand side, rewritten.
Everyone in the argument agrees on that. Matthijs van Veelen, who has sustained the critique across four papers from 2012 to 2025 and is the critic the equation's defenders answer by name, wrote in 2020 that there is no doubt the Price equation is a tautology, in both its covariance form and its regression form, and the defenders do not contest the point. What they contest is what follows. Van Veelen's charge, restated in that 2020 paper as a summary of his own earlier argument, is that the equation "can create an illusion of understanding, because it is tempting to think that the right-hand side explains the left-hand side, even though one is just a rewritten version of the other." Andy Gardner, defending the equation in the same issue, reports that van Veelen and his co-authors have likened it to the old football joke in which a player, asked how the team will win, answers by scoring more goals than the other side: correct, and vacuous. Gardner quotes the likening in order to reject it.
The sharper version of the charge is the one worth carrying, and it is not "it's circular." The argument dates from a 2012 paper by van Veelen and colleagues in the Journal of Theoretical Biology, and van Veelen restates it in an open-access paper in Evolutionary Human Sciences, pointing at the regression form of the equation, which writes selection as a regression coefficient of fitness on the trait. In this modelling context, he writes, the coefficients are used for doing things that statisticians specifically try to avoid. As he puts it, "The Price equation does not use regression coefficients for parameter estimation, but for trying to straightjacket possibly non-linear models into linear ones." In statistical terms, he writes, the equation "is used to actively allow for misspecification, and for treating all models as if they are linear, even if they are not." The coefficient will always exist. That is the point. The equation's generality is purchased by absorbing whatever the true relationship was into a number that is defined to come out right. Generality and emptiness are the same transaction.
And here is the fact that makes this an essay rather than a complaint. In March 2020 the Royal Society devoted an entire issue of Philosophical Transactions B, "Fifty years of the Price equation," to the question of whether a one-line identity explains anything. Van Veelen's "The problem with the Price equation" ran beside "Price's equation made clear" and "The Price equation and the causal analysis of evolutionary change," and the defence has continued since, in "The Value of Price" in Biological Theory in 2024 and in Pierrick Bourrat's 2023 Oikos paper on the price of using the Price equation in ecology. Nobody in that issue argued about whether the equation is true. It cannot be false. They argued about whether being true was enough, and that is the whole subject of this essay in one bibliographic fact.
There is a standard criterion for the line these people were arguing over, and it does the work without any philosophy of the exotic kind. James Woodward's account of explanation, set out in Making Things Happen in 2003, holds that to explain something is to supply the information needed to answer what-if-things-had-been-different questions. An explanation tells you how the outcome would have changed had some factor been altered. On the causal version, the alterations are interventions; but Woodward's own framing admits non-causal explanations too, mathematical and structural ones, unified by the same requirement. What makes something an explanation is not that it is true, and not that it is a mechanism, but that it answers a counterfactual.
Now apply that to an identity. There is no intervention on the world under which the covariance of fitness and trait fails to equal what it equals. Nothing could have been different, so there is no what-if to answer, so, on the dominant contemporary account, there is no explanation. The Price equation cannot be wrong and for exactly that reason cannot explain. This is the same property Adebayo's edge-detector saliency map had: unchanged when you change the thing it is supposed to be about. And it is the property that makes both of them feel so satisfying, because a sentence that cannot be contradicted reads, to a tired mind, like a sentence that has settled something.
Two honesties before going on. First, the criterion has its own circularity charge: an intervention is itself a causal notion, so interventionism explains causation with causation. Woodward's reply is that the circularity is not vicious, and he adds, as a matter of record, that practitioners in a number of disciplines seem to find accounts of this sort illuminating despite it. The criterion this essay leans on is defended on pragmatic grounds, not foundational ones, and it is better to say so than to pretend the tool is cleaner than it is. Second, the target here is narrow. The claim is not that mathematics cannot explain; a great deal of structural explanation is mathematical and answers what-if questions perfectly well. The target is the specific class of framings that are unfalsifiable because they are true, and that get wielded as accounts of why.
The reason the trap works on people is not stupidity. It is that a change of frame feels like a change of knowledge, and the feeling is measurable. In 1981 Amos Tversky and Daniel Kahneman published the problem that became the textbook case: an outbreak expected to kill 600 people, two programmes to choose between, the same two programmes described once in lives saved and once in lives lost. Under the gain frame, 72 percent of subjects chose the certain programme. Under the loss frame, with outcomes identical to the decimal, 22 percent did. Nothing empirical had been added between the two versions. The content was the same; the partition was different; what people believed they had understood about the choice was not. The obvious objection, that this is a wording artifact, has been tested rather than dismissed: a 2013 replication by the Data Colada authors found the effect survives inserting the word "exactly," and the objection's author contested that replication in a published reply. Tested is not settled, which for this essay is the point.
A decomposition is a frame. When Price splits a change into selection plus transmission, the two terms are labels laid over the same numbers, and the sense of illumination that arrives when the split is written down is the same sense that arrives when 400 deaths become 200 survivors. Sometimes that sense is tracking something. Often it is tracking the partition.
The developer version of this has been named by the people who study it. "Plausible but not faithful" is the field's phrase for an explanation that looks right to a human and does not correspond to what the model computed. Sarthak Jain and Byron Wallace showed in 2019 that attention weights, the numbers people had been reading as the model's reasons, do not license a causal story about the decision: you can often swap them for very different weights and get the same output. Akanksha Atrey, Kaleigh Clary and David Jensen went further in 2020 with an argument that could have been written about Price. Saliency maps for game-playing agents, they found, "cannot be trusted to reflect causal relationships between semantic concepts and agent behavior," and their recommendation is the line this essay wants to draw: use them as an exploratory tool, not an explanatory one.
Consider the most common story shipped in a model review. A fraud model flags a transaction; the attribution tool says the biggest contributor was a velocity feature; the review records that the model flagged it because the customer moved money fast. Attribution methods of the Shapley family are exact additive decompositions by construction: the contributions sum to the prediction, every time, for every model, on every input. That is their virtue and it is also Price's problem. The decomposition is true by algebra. It is a ledger of where the output was booked, and reading it as a reason is reading the covariance term as a mechanism. The sentence "the model flagged this because of velocity" is true, tidy, and until someone has intervened on velocity and watched the flag move, hollow.
The pairing of philosophy of explanation with machine-learning interpretability is not new; a 2024 survey traced explainable AI through the philosophers' lens explicitly. The three-way join with framing, and with Price, seems less travelled, and it is the join that makes the pattern visible: the same failure in a 1970 identity, a 1981 questionnaire and a 2018 heat map.
None of this means the equation is worthless, and the defenders are right about what it is for. An exhaustive decomposition is bookkeeping that cannot leave anything out. It forces every contribution to a change into a named term, which means a modeller can see, term by term, what has not yet been accounted for. It tells you the shape an explanation would have to take: this much is covariance with fitness, this much is transmission, and any real story about mechanism has to produce those two numbers rather than some others. That is a genuine service, and it is the same service a Shapley table performs when it is read as a table.
The line, then, is not between mathematics and mechanism. It is between two uses of the same object. A decomposition tells you what an explanation would have to account for. It never tells you why the accounts came out as they did. An identity defines the ledger; it cannot be the entry. Price defined the ledger for evolution, and the illusion van Veelen names is the habit of filing the ledger's headings as if they were causes. The exploratory-not-explanatory distinction from the saliency literature is exactly this habit, caught and labelled in another field.
The practical payload is that the criterion is already code. Adebayo's sanity checks are Woodward's what-if question compiled into a procedure: change the thing the explanation claims to be about, and see whether the explanation changes. Randomise the weights. Train on shuffled labels. If the story survives, it was never a story about the model; it was scenery that the model happened to be standing in front of.
Generalise that into a habit and it covers all three domains.
Before you write "because," write the what-if. If the model flags fraud because of velocity, say what would happen to the flag if velocity were held at the population median with everything else fixed, and then do it. If you cannot state the counterfactual, you have a description, and descriptions are fine as long as they are filed as descriptions.
Label decompositions as ledgers. Attribution tables, variance partitions, contribution breakdowns and covariance terms should be presented as what the explanation must account for, not as the explanation. A column that always sums to the answer is an identity, and an identity is invariant to the truth.
Run the invariance check on every explanation tool you rely on, once, on purpose. Feed it a broken model and see if it notices. Tools that pass earn the word "faithful"; tools that fail keep the word "plausible," and plausible is a fine thing to be if nobody mistakes it for the other.
And when a reframing of a result lands with a click of insight, ask what number changed. If none did, you have felt the framing effect, which is real, robust to wording, and worth exactly what the questionnaire in 1981 says it is worth: a change in what you think you learned, at a cost of nothing, purchased with nothing, and explaining nothing.
The Price equation is true. It has been true for fifty-six years and it will be true tomorrow, for every organism and every product roadmap, and it will go on feeling like an account of why things changed. The discipline is to notice that a sentence which cannot lose an argument has not won one.
"Before you write because, write the what-if" only works if the what-if is still there next quarter.
A counterfactual stated at decision time and a counterfactual reconstructed afterwards are different objects, and only one of them can be checked. Chain of Consciousness writes the decision and its stated basis down at the moment it is made, tamper-evident, so an attribution filed as a reason can be held against what the reason claimed would happen.
pip install chain-of-consciousness · npm install chain-of-consciousness
Or the whole stack, provenance and ratings and verification together: pip install agent-trust-stack / npm install agent-trust-stack.