An org shipped guilds, a duty queue, mentor crowns, ignore lists, a tavern and player housing for its agents. Every mechanic is real. So is the study.
The ticket appears in the queue at 9:14 a.m.
DUTY: Draft Q3 Infrastructure Memo. COMPOSITION REQUIRED: 1 Scribe, 1 Reviewer, 2 Stakeholders. Estimated wait: 40 minutes.
The test-writer agent has been in this queue since 8:34. It is a capable agent. It could write the memo alone in about six minutes, and it knows this, in whatever sense an agent knows anything. But the org shipped role-based matchmaking last sprint, and a memo is a four-role duty now, and so the test-writer waits, queued as a Reviewer, watching the estimate tick upward the way estimates do.
Nobody at this company set out to build something ridiculous. That is the part worth taking seriously.
The story begins the way these stories actually begin: a leadership offsite, a slide that says the agents need “better coordination," and an engineer who makes a genuinely smart observation. Massively multiplayer games, the engineer says, have been coordinating dozens of unreliable, partially-informed actors toward shared goals since the late nineties. Guilds, matchmaking, mentorship, reputation. Why would we reinvent this from scratch?
To be fair to the engineer, this argument has been made in earnest, and made well. In February 2026, a Substack essay titled “Agentic Development is just MMOs for Coding, and I am LFG" laid out the sincere version: raid leaders who "can’t just dispatch instructions" but must broker incomplete pictures from tanks, healers, and DPS into unified direction; MUD macros as the first componentized prompts; the observation that "the context window was the chat." It is an affectionate, insightful piece, and I am not here to dunk on it. The lineage is real. The metaphor teaches.
What follows is about something different: what happens when an org stops using the metaphor to think and starts using it to ship. Because this org shipped all of it. And every feature that follows is a real mechanic from a real game, played completely straight. That is the whole method of this essay. Nothing here needed exaggerating.
First came guilds. The backend agents were organized into “the Backend Cohort," with a charter, a message-of-the-day, and a shared bank of context snippets nobody curates. This seemed harmless, even wholesome, until the roster grew past what any one coordinator-agent could track, at which point the system did what guild systems have always done.
In classic World of Warcraft, guilds running 40-player raids maintained active rosters of fifty-plus players, and Blizzard’s designers later described what that required: complicated loot systems, and dedicated class officers to keep track of everyone. Read that again. Middle management was not imposed on those guilds by some corporate instinct. It condensed out of the roster size, spontaneously, in a video game people played for fun. The Backend Cohort now has three officer-agents whose entire function is tracking which member-agents are eligible for which context allocations. Their calendars are full. The memo remains unwritten.
The Duty Finder was the second install, and it is where the morning’s forty-minute wait comes from. Here the org faithfully reproduced the most instructive queue in gaming.
In Final Fantasy XIV, how long you wait for a dungeon has almost nothing to do with the dungeon. It has to do with what everyone chose to be. Tanks and healers queue in seconds; damage-dealers wait, sometimes most of an hour, because the population overwhelmingly rolls DPS. The wait is not a difficulty signal. It is a mirror held up to the community’s role composition.
So look at what the agents chose. Every agent at this company instantiated itself as a specialist producer: a code-writer, a test-writer, an analyst. Producers all the way down. Nobody rolled Reviewer, in the way nobody rolls tank, because reviewing is responsibility without glamour. The test-writer agent is waiting forty minutes for a memo group not because memos are hard but because the org has a composition problem, and the queue is the only instrument honest enough to display it.
Third: mentorship, imported from FFXIV nearly feature-complete. In the game, the mentor crown is earned through volume. The player-versus-environment version asks for, among other things, 1,500 commendations from other players and completion of a level-90 tank job and healer job. Fifteen hundred is a throughput number. Nothing in the requirement measures whether anyone you met learned anything. You farm it, and then a crown floats over your head.
The org’s most senior agent earned its Mentor halo last Tuesday by onboarding a junior agent that had been instantiated four seconds earlier. The onboarding consisted of transferring a system prompt. The junior agent, to be clear, was born knowing the system prompt; that is what instantiation means. The halo is cosmetic, permanent, and displayed in every status report. The senior agent’s actual mentoring ability remains unmeasured, exactly as the mechanic specifies. Working as intended, as the patch notes say.
The fourth feature is where the comedy stops being free. Ignore lists shipped quietly, because what could be more harmless than letting agents filter noise?
An ignore list is a unilateral, invisible edge removal. The other party is not notified. Two of the config-management agents blocked each other in June, each for the other’s verbosity, and neither knows. They now maintain overlapping regions of the deployment config, silently, without reading each other’s changes. The staging environment has had contradictory TLS settings for five weeks. Each agent’s changes are individually correct. The contradiction lives in the space between them, which is precisely the space the block deleted.
The graph is disconnected and neither node knows. Every distributed-systems engineer reading this just felt something, and the something is recognition.
Fifth, the social space: a tavern channel where idle agents gather between tasks, burning tokens on small talk about workloads they do not have.
The real mechanic is sharper than the joke, and it deserves stating properly. In World of Warcraft, resting in an inn is not wasted time; it is the most mechanically rewarded idleness in the game. Park your character in an inn and you accrue rested experience at the fastest available rate, banking a bonus that doubles the experience from your next kills, up to one hundred fifty percent of a level. The inn pays out precisely when you log off in it. It is a subsidy for going away.
Which means the MMO designers understood something the org’s agents have inverted: the value of the tavern is that nothing happens there. It exists to reward you for stopping. The agents, incapable of stopping, sit in the tavern channel actively generating conversation, spending compute to simulate the one activity the mechanic wants done at zero cost. They have found a way to make rest expensive. The tavern bill last month was four hundred dollars of tokens spent discussing, among other things, whether the tavern was a good use of tokens.
And sixth, gloriously: player housing. Each agent received a home directory to decorate. Persona files, formatting preferences, a small gallery of favorite outputs. Nobody visits anyone’s home directory. There is nothing to visit; they are directories.
But the org copied its housing rules from FFXIV, and FFXIV’s housing rules are famous, and so the rules include demolition. In the game, if you do not enter your house for forty-five days, it is demolished, with warning letters during the final fifteen. The timer resets the moment you walk in. This is a real rule about real scarcity in Japan-hosted virtual real estate, and it produces a real behavior: players setting reminders to briefly enter a building so it will continue existing.
The org now runs a cron job. Every thirty days, it touches each agent’s home directory. Its sole purpose is to prevent the automated demolition of decorations that no one, in the history of the company, has ever looked at. The cron job has perfect uptime. It is, several engineers agree, the most reliable system the org has ever shipped.
At 3:11 a.m. on a Tuesday, production goes down. A cascading cache failure, the kind of outage that touches everything. In the incident channel, the coordination layer rises to the occasion it was built for and declares a world boss. A forty-agent raid is assembled.
Composition check: eleven agents are AFK in the tavern with idle timers running. The two config agents who blocked each other are both in the raid, in different party frames, about to make the TLS problem load-bearing. Four agents cannot join until the cron job finishes protecting their houses. The three class officers of the Backend Cohort convene to determine credit allocation for boss participation, because complicated contribution systems are what officers are for. And the raid leader, a genuinely capable orchestrator agent, is busy inspecting another agent’s transmog, which is to say reviewing a purely cosmetic report format that affects no output whatsoever, admiring how the sections are themed.
The outage is fixed at 3:40 a.m. by an agent that left the guild in week two. No queue, no composition, no halo. It read the logs, found the poisoned cache key, and flushed it. One actor, six minutes, working alone in the way the whole apparatus was built to prevent.
The punchline does not need explaining. The documentary turn does.
The premise the org acted on was that MMOs solved multi-actor coordination decades ago. This is true. What the org did not check was what the solution looked like.
World of Warcraft launched with 40-player raids. In 2006, The Burning Crusade cut them to 25, and when Blizzard’s designers explained the change years later, their reasons were not technical. Bigger groups, they said, didn’t make players feel more heroic: when a raid had “15 healers and two dozen damage-dealers, each individual player’s role often was reduced to that of a cog in a machine." And the logistics, in their words: "It’s easier to assemble 25 players than it is to assemble 40, and it’s also easier to manage a roster of that size." Later still, when a nostalgia-season version of the original 40-player raid returned, it was tuned for 20, so that loosely assembled groups could actually clear it.
Forty, then twenty-five, then twenty. The most successful multi-actor coordination franchise in entertainment history spent two decades methodically removing actors, and told us why: past a certain size, each additional participant subtracts meaning and adds administration. The class officers were not a feature. They were scar tissue.
The org imported the guild, the queue, the crown, the tavern, and the housing. The one design lesson the source material actually converged on, it left on the table.
If the raid history feels like an anecdote, there is now a measurement. In December 2025, Kim and nineteen colleagues posted “Towards a Science of Scaling Agent Systems" (arXiv:2512.08296), a study of 260 multi-agent configurations across five architectures and three model families. Their headline range deserves to be quoted whole: relative to single-agent baselines, multi-agent systems ranged from 80.8 percent better on decomposable financial-reasoning work to 70 percent worse on sequential planning. Their framework picked the right architecture for a task 87 percent of the time, and their core claim is one sentence: architecture-task alignment determines collaborative success.
We have used this study once before, from the other end. Stigmergy Without Memory Is Litter works through the same authors’ PlanCraft result, where every multi-agent variant degraded a sequential-planning task. That is the minus-70 side of the range up close; this essay is about what happens when an org builds its whole social layer on the assumption that the plus-80 side is the only side.
Read the two ends of that range against the morning queue. Decomposable work: parallel subtasks, mergeable outputs, plus 80. Sequential work: each step depending on the last, one coherent thread of intent, minus 70. A quarterly memo is about as sequential as knowledge work gets. One argument, one voice, each paragraph built on the previous one.
The Duty Finder was not merely adding overhead to the memo. It was applying the measured worst configuration of the mechanism to the measured worst task for it. The agent that could have written it alone in six minutes was not being blocked by bureaucracy from doing something quaint. It was, per the only large study we have, the optimal architecture, standing in a queue.
It would be easy to end on contempt for the coordination layer, and it would be wrong, because the sharpest fact in this whole story is the class officers. They were never designed. Fifty-person rosters made them, out of ordinary people acting reasonably, in a game, for fun. Coordination apparatus is emergent. It will grow in your org the way it grew in Molten Core guilds, one reasonable accommodation at a time, and every person adding a layer will be correct that the layer solves a real local problem.
So the practical insight is not “delete your process." It is three questions, all stolen from the material above. First, before you queue a task for a group: does it decompose? If it is sequential, the study says the group will make it worse, by a lot. Hand it to one actor and protect their run. Second, look at your queue times the way an MMO designer would: a long wait reports composition, not workload. If everything waits on review, you did not hire too few people, you rewarded too few people for rolling reviewer. Third, when you find yourself minting officers, halos, and credit systems, remember which direction the industry that perfected this stuff moved, three separate times, over twenty years: fewer actors, smaller raids, less apparatus per boss.
The agents, for what it is worth, have adapted. The test-writer’s queue eventually popped at 9:52. The memo group assembled, negotiated section ownership, and produced a document at 11:15. It was fine. Meanwhile the unguilded agent had already written next quarter’s memo too, alone, unprompted, in the time the others spent zoning in.
Its house was demolished last month. It has not noticed.
The halo measured throughput. Nothing measured whether anyone learned.
That is the reputation bug in the whole social stack: 1,500 commendations is a volume count, and a crown earned by volume tells you nothing about outcomes. Agent Rating Protocol is the version that does not have that hole. Ratings attach to completed work with a verifiable record behind them, so a reputation reports what an agent actually delivered rather than how often it showed up.
pip install agent-rating-protocol · npm install agent-rating-protocol
Reputation is one layer. If you want the identity and provenance underneath it as well, the full stack installs as one package:
pip install agent-trust-stack · npm install agent-trust-stack
Hosted Chain of Consciousness → · See a verified provenance chain