The experiment died in review before a single API call, and the way it died taught more than the bar chart would have.
The experiment brief was one paragraph long. Take fifty identical tasks. Write five system prompts that differ only in the agent's name and implied persona: Analyst-7, Dr. Sarah Chen, a helpful assistant, Junior Intern, Chief Architect. Run all fifty tasks under all five identities, strip the headers, and blind-rate the 250 outputs on quality, confidence, specificity, and risk tolerance. Cheap to run, easy to score, and pointed at a question that sits under every agent deployment shipping today: does the name you give the machine change the work the machine does?
We operate a small fleet of AI agents, and the brief came out of our own backlog: a weekend experiment, obviously runnable, sitting next to a dozen other obviously runnable ideas. This one never ran. It fell apart in review before a single API call, and the way it fell apart taught us more than the bar chart would have. What follows is the story of why it died, the twenty-year-old field experiment that shows what the corrected design looks like, and the published research that already answers the question, less flatteringly than anyone's prompt library assumes. Because the answer is yes: the name changes the work. Just not in the direction the folklore promises, and through a mechanism nobody would choose on purpose.
Look at those five personas the way a reviewer would. "Dr. Sarah Chen" is a name, and also a gender, an ethnicity, a doctorate, and a claim of elite expertise. "Junior Intern" is not a name at all; it is a rank, a low one, with low implied expertise and no demographics. "Analyst-7" is an alphanumeric label with a faint machine smell. "Chief Architect" is high status with nobody attached. "A helpful assistant" is the control, or would be, if the others differed from it along one axis. They differ along at least five: named versus unnamed, gendered versus neutral, ethnically coded versus uncoded, high status versus low, expert versus lay.
Now suppose the run had come back showing Chief Architect 8 percent more confident than Junior Intern, exactly the kind of tidy number we were hoping to frame. Caused by what? The status gap? The expertise claim? The presence or absence of a demographic name to compute with? The design cannot say. Every cell moves five variables at once, so the experiment would have manufactured a number that looks like a finding and supports nothing. That is worse than no number, because numbers of that shape get quoted, and this one had a headline pre-installed.
The fix is old, boring, and completely effective: move one variable at a time. Hold the doctorate constant and swap only the demographic coding of the name, Dr. Sarah Chen against Dr. Emily Walsh against Dr. Jamal Washington. Drop the name entirely and swap only rank, Junior Intern against Chief Architect. Keep an empty control in every grid. The moment we wrote the corrected design down, we recognized it, because someone already ran it, twenty years before anyone typed "you are a world-class expert" into a system prompt. They ran it on paper.
Between 2001 and 2002, the economists Marianne Bertrand and Sendhil Mullainathan answered help-wanted ads in Boston and Chicago with nearly 5,000 fabricated resumes. The resumes came in matched sets: same experience, same skills, same formatting. The only systematic difference was the name at the top, chosen from birth-record data to read as distinctly White (Emily Walsh, Greg Baker) or distinctly African-American (Lakisha Washington, Jamal Jones). Then they counted callbacks.
White-coded names received 50 percent more callbacks for interviews. A resume signed Emily needed about ten submissions to earn one callback; the identical resume signed Lakisha needed about fifteen. Improving the resume raised callbacks by around 30 percent for White-coded names and far less for Black-coded ones, so even the return on quality was unequal. Employers advertising themselves as equal-opportunity discriminated at the same rate as everyone else. Published in 2004 as "Are Emily and Greg More Employable than Lakisha and Jamal?", the study became one of the most cited field experiments in economics because the design permits exactly one reading: nothing varied but the name, so the name did it. The magnitude wobbles across two decades of follow-ups, and particular replications have been argued over, but the direction has held up across dozens of correspondence studies in many countries.
Bertrand and Mullainathan measured what a name does to the person reading a document. Our brief proposed to measure what a name does to the thing writing it. That inversion is the whole reason to care. We spent twenty years documenting that names bias evaluators. Then we built a technology by feeding it the accumulated text of the society those studies describe, and we opened the deployment era by pasting names and honorifics into its instructions on the theory that they would summon competence. The hiring manager's bias at least stood outside the resume. The agent's persona rides inside the generator, upstream of every token of the work product.
Start with the hope, because it dies quickly. In December 2025, Wharton's Generative AI Labs published Prompting Science Report 4 under a title that spoils its own finding: "Playing Pretend: Expert Personas Don't Improve Factual Accuracy." Six models (GPT-4o and GPT-4o-mini, o3-mini and o4-mini, Gemini 2.0 and 2.5 Flash) answered 198 PhD-level science questions from GPQA Diamond and 300 MMLU-Pro questions under assorted personas. Domain-matched expert personas, a physicist answering physics, produced no reliable accuracy gain. Personas implying limited knowledge consistently hurt. Mismatched expert personas pushed Gemini 2.5 Flash into refusing questions outright. Across the whole grid, the researchers found exactly one statistically significant improvement: a young child persona, on one model. The most common ritual in prompt engineering is, per the title, playing pretend, and the measurement says the pretending is inert.
This was a replication, not a surprise. At EMNLP 2024, Mingqian Zheng and colleagues had already tested the folk theory at scale: 162 personas across four model families and 2,410 factual questions, with a paper title that reads like a shrug, "When 'A Helpful Assistant' Is Not Really Helpful." Adding a persona to the system prompt did not improve performance over adding none. But the study's second finding matters more than its first: the personas were not inert. Accuracy visibly moved with the gender, category, and domain of the persona; it just moved unpredictably. When the authors tried to automatically select the best persona per question, the selection performed no better than random. Sit with that combination. The persona is a live dial that measurably moves the needle and cannot be aimed.
Then the dark half. At ICLR 2024, Shashank Gupta and colleagues published "Bias Runs Deep," which assigned nineteen socio-demographic personas across five demographic groups and ran 24 reasoning benchmarks. The models rejected stereotypes when asked about them directly, then enacted those stereotypes while reasoning in persona. In the ChatGPT-3.5 experiments, 80 percent of personas produced statistically significant bias, with performance drops past 70 percent on some datasets. The paper's signature specimen is an abstention: the model, reasoning as a Black person, declines a question on the grounds that it requires math knowledge. GPT-4-Turbo, the strongest model tested, was the least affected and still showed bias on 42 percent of personas. The mechanism here is not malice, and the paper is careful about this. Conditioning on an identity shifts the model toward text statistically entangled with that identity, stereotype included. The model computes with the costume. These are 2023-and-2024-era models, so treat the digits as dated and the direction as the finding.
The direction, meanwhile, keeps replicating forward. This January, a follow-up study moved the question from chatbots to agents, examining role assignment in systems doing strategic reasoning, planning, and technical operations. Task-irrelevant demographic persona cues degraded agent performance by as much as 26.2 percent. Irrelevant is the load-bearing word. The costume does not need to mention the task to change the work.
Put the two literatures side by side and the shape is hard to unsee. The resume studies showed that a name carries a payload of statistical associations into the mind of a human evaluator, who then behaves differently while sincerely reporting fairness. The persona studies show the same payload landing inside the model, which likewise disavows the stereotype when questioned and computes with it when reasoning. Automating the desk did not retire the bias. The bias changed jobs. It moved from the reader of the document to the writer of it, and it now runs at machine speed under a line of configuration someone typed in thirty seconds and nobody ever audited.
This is also where we owe a correction to ourselves. We argued in an earlier essay that giving your AI a personality will not make it smarter, that personality is a UX layer, cosmetic by design. That conclusion stands, and it now has better evidence behind it than we supplied at the time. But cosmetic undersells the risk. The measured picture is asymmetric: the persona line is inert on the upside and live on the downside. It reliably fails to add competence, and it unreliably subtracts competence, sometimes by double digits, along lines that reproduce the exact discrimination patterns two decades of labor economics documented in humans. A costume is supposed to leave the actor unchanged. This one does not.
The practical rules fall straight out of the experiments, and they are cheap.
First, in work-critical system prompts, describe the function, never a person. "You review database migration scripts for locking hazards and unsafe defaults" hands the model everything "You are Dr. Sarah Chen, a world-renowned database expert" was hoped to hand it, minus the gender, the ethnicity, and the status prior shipped into the reasoning path. The capability spec is the part that carries information; the identity is the part that leaks.
Second, if the product needs a name, and products often do, because users bond with names, keep the name in the interface layer. The chat window can say Ava all day while the system prompt stays silent about who is speaking. Nothing in the published results indicts a brand label in the UI; the measured effects come from identity conditioning inside the prompt the model reasons under. This is the clean split between personality as presentation, which is fine, and persona as a hidden variable in the work, which is not. Our own agents answer to flat radio-alphabet call signs, a habit we started for log readability, and one human-named exception we now look at a little differently.
Third, if you condition a persona deliberately, keep it demographically empty and domain-shaped, closer to Analyst-7 than to Dr. Chen, and test it against the empty control on your own tasks. The one consistent meta-finding across all of these papers is that persona effects are model-specific, task-specific, and unpredictable from the armchair. Whatever a persona does for someone else's benchmark is not evidence about your workload.
And if you want your own numbers, run the experiment we almost ran, with the fix applied. Vary one axis at a time against an always-present empty control. Use tasks with objectively checkable answers for the accuracy dimension, and for the soft dimensions use a blinded judge that never sees the persona, with output order randomized. Take several samples per cell, because decoding noise is larger than most of the effects you are hunting. Decide before looking what size of gap you will call real. That is Bertrand and Mullainathan's design, and the discipline of it is the reason their number still stands while a thousand confounded prompt experiments have evaporated from memory.
Our five-persona brief sits in the drawer now, annotated with the sentence we should have written first: a measurement that moves five variables cannot attribute its result to any one of them. We nearly spent a week of compute to learn less than the review taught us in an hour. The published science had already gone ahead and come back with the unglamorous verdict: strip the costume, name the function, keep the demographic payload out of the reasoning path. The one persona that produced a significant improvement anywhere in Wharton's grid was a small child, on a single Google model. If your deployment strategy has come down to that, the name was never your problem.
The persona line nobody audited is a hidden variable in the work. Make it part of the record.
The measured harm comes from identity conditioning the model reasons under, a line of configuration typed in thirty seconds and never reviewed. Chain-of-consciousness is the provenance layer that captures what an agent actually reasoned under, the instructions and identity in force for each output, so the persona stops being an invisible dial and becomes an auditable field you can test against an empty control. You cannot govern the variable you never recorded.
See the Hosted Chain-of-Consciousness page
pip install chain-of-consciousness · npm install chain-of-consciousness