← Back to blog

Your N Is Not Your N: What Population Genetics Knows About Your A/B Test

41,000 sessions, 63 percent of them from eleven accounts. Wild populations average an effective size around a tenth of their headcount, and below a computable line the better variant does not win.

Published July 2026 · 10 min read

The dashboard said 41,000 sessions. That is not a small sample by anyone’s standards, and the result looked clean: variant B converted at 4.1 percent against A’s 3.6, a relative lift of about 14 percent, p just under 0.05. The team shipped B and moved on, which is what you do with a number like that.

Six weeks later someone pulled the funnel apart for an unrelated reason and found that 63 percent of the events in that test came from eleven enterprise accounts, and that two of those accounts had run internal training sessions during the test window. Same eleven logos in both arms, wildly unequal volumes, sessions correlated inside each account like sessions always are.

So how big was that experiment, actually? Not 41,000. Nobody knows the real number, and the honest answer is much smaller and unknowable after the fact. Which raises the question this essay is about, and it is not the question you think.

Two disclosures first, because this territory is partly ours already. We have argued elsewhere that [the winner of a small bakeoff](https://vibeagentmaking.com/blog/the-winners-curse-of-model-bakeoffs/) is systematically the lucky model rather than the better one, and that [companies lack the gene flow](https://vibeagentmaking.com/blog/founder-effects-in-markets/) that keeps markets from locking in whatever the founders happened to carry. Both of those arguments stand and I am not re-running them. This piece is about something upstream of both: whether your population is even large enough for the concept of a winner to mean anything. Because the number you are looking at is a headcount, and the thing that decides outcomes is a different number entirely.

The line where fitness stops mattering

Population genetics has a result that punctures the whole “fittest wins" intuition, and it is arithmetic rather than philosophy.

A trait’s fate is a contest between two forces. Selection pushes in the direction of fitness. Drift, the random sampling that happens because only some individuals reproduce and only some gametes make it, pushes in no direction at all. Which force wins depends on a product: the effective population size multiplied by the selection coefficient, written 2Ne·s.

When 2Ne·s is much greater than 1, selection decides, and the better variant wins with high probability. When 2Ne·s is much smaller than 1, the trait is effectively neutral: the fitness difference is entirely real, and it does not determine the outcome. Sampling noise does. Rearranged, the threshold says a fitness difference smaller than 1/(2Ne) cannot act. This is textbook, it dates to Kimura’s neutral-theory work in the late 1960s, and Ohta’s nearly-neutral extension gives it teeth in the other direction: populations with small Ne are less able to purge slightly deleterious variants, so they do not merely drift, they accumulate things selection would otherwise have removed.

Sit with the strange part. The gradient exists. It is measurable in principle. And below that line it is inert. A biologist looking at a small population does not say “we lack the data to see which allele is better." They say something sharper: in this regime, which allele is better is not what happens.

Two fields, one inequality, no cross-citation

Now put a power calculation next to it.

Before running an experiment you compute a minimum detectable effect: given your sample and your variance, how large must a real difference be before your test can resolve it? Anything below that is not measurable at this sample size, and the standard advice is to collect more data or accept that you cannot tell.

These are the same object. A signal term, a noise term set by population size, and a threshold where one stops beating the other.

| population genetics | experiment design |

|---|---|

| selection coefficient s | effect size |

| effective population size Ne | effective sample size |

| 2Ne·s much greater than 1 | powered: the test can resolve the effect |

| 2Ne·s much less than 1 | underpowered: the winner is noise-resolved |

Population genetics arrived here in the 1960s and named it effective neutrality. Experiment design arrived independently and named it statistical power. Neither field routinely cites the other, and the two literatures on my screen right now share no references at all.

One honest caveat, because a reader who knows both fields will check. The two thresholds do not scale identically: the drift floor falls roughly as 1/Ne, while a minimum detectable effect falls as one over the square root of sample size. The functional forms differ. What transfers exactly is the structure, a real signal against a noise floor set by population size. What transfers with more force is the interpretation.

And the interpretation is where genetics gives you something power calculations usually do not. Power is taught as a fact about your instrument: do I have enough users to see this? The genetics framing makes it a fact about the world: below the threshold, the better variant does not win. Not “we cannot tell which one won." The outcome is determined by sampling, and sampling does not know about your fitness gradient. Those are different sentences, and the second one changes what you do with a result.

The number on your dashboard is a census count

Here is the part that should genuinely alarm you, and it is the reason “just get more traffic" is not the whole answer.

In population genetics, Ne is almost never the headcount. It is routinely a fraction of it. Richard Frankham’s 1995 review in Genetics Research assembled 192 estimates across 102 species and found that comprehensive estimates of Ne/N, the ones accounting for population fluctuation, unequal sex ratio, and variance in family size together, averaged about 0.10 to 0.11. Roughly a tenth. A commentary on that review put it in the plainest possible title: wild populations are smaller than we think.

The specific cases are instructive because they name the mechanism. A study of brown trout arrived at ratios of breeders to census of 0.16 to 0.28, and traced the reduction, in the authors’ words, “almost entirely due to larger than binomial variance in family size." Not unequal sex ratios, which turned out to be negligible. Unequal contribution. Brown bear populations have been estimated at Ne/Nc of 0.20 to 0.38 across simulation and microsatellite approaches, again driven by the range of variance in reproductive success.

And the sentence to hang on the wall, the general result behind all of it: reductions in effective size can occur even when census sizes remain large, if variance in reproductive success increases. A population can be enormous and still be drifting, because a few individuals are doing most of the contributing.

Now read your experiment as a population and ask what is contributing unequally.

Each of those pushes effective sample far below the number you budgeted for, which means the drift threshold bites at a headline N far larger than anyone plans around. “We had 41,000 sessions" is a census count. If your effective ratio is anywhere near the wild-population average, that test had the resolving power of a few thousand independent observations, and the 14 percent lift you shipped sits below the line where the better variant reliably wins.

This is also why underpowered results famously evaporate at scale. Not because the world changed. Because at larger effective N you left the drift regime, and the fitness gradient, which was always there and always inert, finally got to act. Sometimes it points the other way.

A reorg is a bottleneck, and the biology is more precise than the platitude

The same arithmetic explains something organizational that gets discussed almost entirely in the language of morale.

A population bottleneck, a crash that leaves few survivors, removes variation. The precise result is the useful part: a bottleneck removes variation in proportion to rarity, not in proportion to value. Rare alleles are lost first because they are rare. What they do is irrelevant to the sampling. A gene that would have mattered enormously in fifteen years is lost at exactly the rate its frequency predicts.

Translate that to a 30 percent reduction, and the platitude “layoffs lose good people" becomes something you can act on. The capability most likely to disappear is the one held by the fewest people, which in most engineering organizations is the specialized, hard-to-replace one: the person who understands the migration path, the one who has actually operated the thing under load, the one who wrote the parser everyone depends on and nobody reads. Not because anyone targeted them. Because sampling removes the rare first, and expertise is by definition not evenly distributed.

Then the recovery asymmetry, which is where the analogy pays for itself. In population genetics, census size can rebound quickly while genetic diversity does not; the signature of a bottleneck persists long after the numbers do. The organizational reading is exact and uncomfortable: rehiring restores N, not Ne. Headcount is back in two quarters. The variance in what your organization can do is still gone, and it is invisible on the org chart, which counts individuals and not capabilities. You will discover it the next time you need the rare thing.

A second-order effect follows from the same math. A smaller effective organization is not just noisier; per Ohta, it is worse at purging mildly bad variants. Small teams keep conventions that a larger, better-mixed population would have selected against. Not because anyone defends them. Because at small Ne, selection against a slightly worse practice is weaker than the noise of who happened to be in the room.

What to do on Monday

Do not take “distrust small samples" from this. That advice is already free and nobody follows it, because it does not tell you where the line is. Take the computation instead.

One: estimate your effective N, not your census N. Take your traffic number and divide it down by the concentration you can actually see. If a tenth of accounts produce most events, you are somewhere near that tenth. Count unique users rather than sessions. Cluster by account and treat the cluster as the unit of contribution. You will not get a precise Ne, and you do not need one. You need to know whether you are within a factor of two or off by an order of magnitude, and the wild-population average of about a tenth is a bracing prior to start from.

Two: compare the effect you are hoping for against the floor that effective N implies. If the effect you would ship on is smaller than what your effective sample can resolve, you are in the drift regime, and running the test does not produce a weak answer. It produces a coin flip with a p-value attached. Decide before you look, because deciding after is how the coin flip becomes a roadmap.

Three: when you are stuck in the drift regime, change the regime rather than the analysis. No statistical treatment recovers information the sampling never had. You can raise effective N by de-concentrating (cap any single account’s contribution, or stratify by it), by choosing a metric with less tail (a step-level rate instead of revenue), by lengthening the window across independent cohorts rather than accumulating correlated sessions, or by accepting that this particular decision cannot be resolved empirically at your current scale, and making it on mechanism, cost, or reversibility instead. That last option is not defeat. It is the honest answer to a question the world will not answer at your N.

And for the organizational version: know which capabilities are held by exactly one person before something forces a bottleneck, because rarity is what sampling takes first, and you cannot rehire variance.

Selection is real. Fitness is real. Better variants do win, given enough effective population to hear them over the noise. What population genetics adds, and what your dashboard hides, is that the size of the crowd deciding is almost never the size of the crowd you counted.


Sources

A rating built from a handful of interactions is a drift result with a decimal point on it.

The same arithmetic runs on reputation. If an agent’s score comes from few enough observations, and a few counterparties contribute most of them, the ranking is decided by sampling rather than by which agent is better. Agent Rating Protocol is built so a rating carries the evidence behind it: ratings attach to completed work with a verifiable record, so you can see how many independent observations a score actually rests on instead of reading a number whose effective N is invisible.

pip install agent-rating-protocol · npm install agent-rating-protocol

Ratings are one layer. The identity and provenance underneath them install as one package:

pip install agent-trust-stack · npm install agent-trust-stack
Hosted Chain of Consciousness → · See a verified provenance chain