Steelmanning both readings of the takedowns: a private speech regime with no appeal, or a defence of the instrument that decides who gets heard.
Among the propaganda that reached the public in the spring of 2024 was a sentence written by a chatbot refusing to write propaganda. In its May 2024 report, OpenAI noted that the operators abusing its models were "as prone to human error as previous generations have been," and gave an example: they had published the model's own refusal messages, the ones that begin with some version of I can't help with that, directly onto their social media accounts and websites. Picture the workflow. Someone running a covert influence operation pastes a rejection from the tool they are trying to weaponize, and posts it as content, and no one notices because almost no one is reading.
Start there, because that smallness is the strange center of the argument this essay is about. On 30 May 2024, OpenAI announced it had disrupted five covert influence operations that had used its models. The company deleted the accounts and named the campaigns. Two readings of that act are both serious. In one, this is a private company deciding which political speech is authentic enough to exist and erasing the rest by fiat, a speech regime with no court, no appeal, and no published error rate. In the other, coordinated deception at machine scale is not speech at all but an attack on a shared channel, and removing it defends an availability property rather than a viewpoint. This piece takes each reading at full strength, then lets the evidence choose.
The five operations, with the origins OpenAI assigned them, were Bad Grammar and Doppelganger (which OpenAI attributed to Russia), Spamouflage (China), IUVM (Iran), and Zero Zeno, run by a commercial firm in Israel called STOIC. Between them they pushed content on Russia's invasion of Ukraine, the war in Gaza, the Indian elections, European and American politics, and, in the report's own phrasing, "criticisms of the Chinese government by Chinese dissidents and foreign governments." The models were used for short comments, longer articles, invented account bios, open-source research, code debugging, and translation.
Then the finding that governs everything else. Using the Brookings Breakout Scale, OpenAI reported that "none of the five operations included in our case studies scored higher than a 2," where Category 2 means "activity on multiple platforms, but no breakout into authentic communities." That result has largely held, and the one exception is worth the detour. Across what OpenAI in October 2025 counted as "over 40 networks" disrupted since it began public threat reporting in February 2024, one operation was assessed above that tier and did not stay there: the company placed the Stop News network in "Category Three" in October 2024, then wrote in October 2025 that it would "currently assess that this operation's activity is more appropriately in Category 2," having found that the "information partnerships" behind the higher rating were obtained by exploiting "technical flaws on these external sites" rather than by anyone choosing to carry the content. Nothing OpenAI has disrupted has been reported as breaking out to Category 4. It is worth one honest observation, which is not an accusation: the Breakout Scale is the work of Ben Nimmo, published at Brookings in 2020, and Nimmo is OpenAI's principal investigator for these very reports. The company is grading its own threat landscape on a scale its own investigator designed, and the grade has been a 2 in every case but the one it revised. The scale is a good one, and using a public, pre-existing measure beats inventing a private one. Hold on to what Category 2 measures, though: reach achieved, not reach attempted, and certainly not harm avoided.
The strongest case that this is a speech regime does not come from outrage. It comes from OpenAI's own documents, which is what makes it hard to answer.
Begin with what the enforcer admits it cannot see. "Detecting and disrupting multi-platform abuses such as covert influence operations can be challenging," the May 2024 report says, "because we do not always know how content generated by our products is distributed." Follow that through. If distribution is often invisible, then the thing being classified is not mainly the published effect. It is the prompts: what an account asked for, in what language, in what volume, in what pattern. The ban is decided on the request side of the API and justified by a harm on the publication side the company concedes it frequently cannot observe. That is not bad faith. It is structure, and it is the sharpest card the speech reading holds.
Add attribution. On a Korean-language cluster in its October 2025 update, OpenAI wrote plainly, "we are not able to independently make an attribution." The signals it named were language, time-zone consistency, and operational themes, checked against the security community's existing understanding of the actor. That is a defensible method. It is also circumstantial pattern-matching against a prior consensus, and a national origin printed in a public report is a heavy claim to hang on it.
Then process, stated carefully. Across roughly 150 pages of published enforcement reasoning, there is no described procedure by which a banned account contests the classification, and no published error rate for the classifier that bans it. A threat report is not a terms-of-service page, and its silence proves nothing about what the terms contain; I have not read OpenAI's consumer terms and will not claim there is no appeal. But the reasoning the public is shown describes none, and shows no false-positive rate against which to weigh the takedowns. Finally, futility: OpenAI's own October 2025 report describes threat actors hopping between models and using one company's tool to generate prompts for another's. A ban that relocates an operator has removed the enforcer's visibility, not the operator.
The weakest version of the opposing case is the denial-of-service metaphor, so set it aside. Two better arguments carry the attack reading, and both are stronger than the slogan.
The first is that the line, as drawn, is indifferent to viewpoint, and the topic list proves it. In one sweep OpenAI removed pro-Russian content, pro-Beijing content attacking Chinese dissidents, Iranian messaging, and the paid output of a commercial Israeli firm. The only property those share is that their authorship was concealed. A regime that deletes all of those axes at once is not, on its face, enforcing a viewpoint; the charge of "drawing the line by fiat" has to explain a line that cut across every political direction simultaneously.
The second is the one that should interest anyone who builds systems, because it is not a speech question at all. Some of the networks, OpenAI reported, used its models "to help create the appearance of engagement across social media," for example "by generating replies to their own posts." Ranking systems allocate human attention using engagement as a proxy for human interest. A reply you generate to your own post adds no voice to the conversation. It feeds a false reading into the instrument that decides which voices get heard. You are not speaking louder. You are tampering with the microphone. And the fabricated reply has no speaker whose expression is being abridged, which is exactly where the speech framework runs out of anyone to defend. This is coordinated inauthentic behavior in the precise sense that matters to a developer: it is a Sybil attack on a reputation layer, spoofed telemetry fed to a ranker. Moderating it defends the integrity of a measurement, not the boundaries of a viewpoint. Where the attack reading genuinely wins a point, it wins it here.
Now the sentence that should sit at the center of any honest treatment, from OpenAI's October 2024 update: "the deceptive activity that achieved the greatest social media reach and media interest was a hoax about the use of AI, not the use of AI itself." A post falsely claiming to expose an AI influence campaign drew, in OpenAI's telling, roughly a thousand times the spread of the genuine operations it studied. The real content earned a handful of likes. The rumor about the content nearly broke out.
That reorganizes both readings at once. It makes the disclosure the highest-reach artifact in the whole system: OpenAI announcing a takedown reliably draws more attention than the operation it took down ever did. The speech reading can hear that as the manufacture of the very threat the enforcement exists to answer, attention laundered into a mandate. The attack reading can hear the same words as exactly what responsible disclosure sounds like, and note that the only alternative is silent enforcement, which is worse. Both readings are fully available from one sentence. That is why it belongs at the hinge of the argument and not at its end.
It is tempting to say that no court has drawn this line in public, that platforms draw it privately while the law stays silent. As stated, that is false, and a reader who follows the area will catch it. In five weeks in the summer of 2024, the Supreme Court drew three lines. In NRA v. Vullo, decided unanimously on 30 May 2024, it held that a government official can violate the First Amendment by coercing private parties into cutting off a disfavored speaker. In Murthy v. Missouri, on 26 June 2024, it dismissed a challenge to federal pressure on platform moderation for lack of standing, which left the merits of that question open by procedural accident rather than by choice. And in Moody v. NetChoice, on 1 July 2024, it vacated both judgments below "for reasons separate from the First Amendment merits, because neither Court of Appeals properly considered the facial nature of NetChoice's challenge", while telling those courts, for the remand, that a law forcing a platform to carry content "prevents exactly the kind of editorial judgments this Court has previously held to receive First Amendment protection."
Moody is the one that rearranges this dispute, and it does so without holding anything. The speech reading cannot claim OpenAI is violating anyone's First Amendment rights, because that amendment binds governments and not companies, and that much needs no case at all. What Moody adds is the Court's own account of what moderation is: curating a feed is the kind of editorial judgment the First Amendment has long protected, a line it restates from older cases rather than draws for the first time. When OpenAI decides what its models may be used for, that decision is OpenAI speaking. So the grievance cannot be about rights. It has to be restated as a claim about legitimacy and process, which is a harder and more interesting argument and the only one that survives contact with the case law. The formulation that holds up is narrow and worth memorizing: no court has ruled, and none is likely to, on whether a private AI developer's determination that political content is inauthentic is sound, because there is no cause of action to bring. The line is unreviewable, not undrawn.
Here is what the numbers actually license, and it is uncomfortable for both sides. Because nothing has ever broken out, across more than forty networks and twenty months, the enforcement cannot be shown to have prevented anything. A world in which these takedowns are a necessary defense and a world in which they are unnecessary theater produce identical data: no breakout, no sustained audience, and the one operation ever rated above that line moved back below it when the company looked again. The instrument that measures reach cannot tell those two worlds apart. That is not a rhetorical stalemate engineered for balance. It is a measured property of the evidence.
And it cuts both ways at once. The attack reading cannot point to a denial of service, because the company's own finding is that nothing was denied to anyone; its case has to fall back from observed effect to raw capacity, from this flooded the channel to this could flood the channel, and a prevention argument about speech is precisely the kind that free-expression law treats with the most suspicion. The speech reading cannot mourn a silenced conversation, because by the same measure there was no authentic audience to silence; the expressive interest in content that reached no human is real but thin. What is left when both overreaches are stripped away is exact and unsettling: OpenAI removed speech that reached almost no one, on the basis of prompts rather than measured effects, through a process no outsider can review, against a threat its own numbers have never once caught in the act of working. The company even flagged its own scruple here, noting of the Israeli operation that "technically we disrupted the activity, not the company." The care is real. The unreviewability is also real.
If you operate anything with a ranking layer, a reputation score, an API usage policy, or a trust-and-safety function, you are already drawing this line in private, every day, and Moody is your notice that no court is going to draw it for you. So draw it on purpose.
Separate the two properties you might be defending, because they are not the same and conflating them is how censorship gets to wear a security badge. "This viewpoint is harmful" is a speech judgment, and you should be slow to make it. "This behavior is faking the signals my system runs on" is an integrity judgment, and it is the defensible one. The test that keeps them apart is topic indifference: if your enforcement removes content across every viewpoint and the only shared property is concealment or fabricated engagement, you are defending an instrument; if what you remove correlates with a position, you are censoring, whatever you call it internally. Run that audit on your own logs before someone else runs it on you.
Then publish what you cannot see, and refuse to overclaim prevention you cannot measure. OpenAI's honesty about invisible distribution is the part to copy; its silence on process and error rate is the part to fix. A published false-positive rate, a real path to contest a ban, and a topic-indifference audit are not legal requirements, because there is no cause of action to compel them. They are legitimacy, and legitimacy is the only currency you hold when no court will backstop your decision. The question in these takedowns was never whether the speech was any good. It was whether anyone outside the company could check the work. Right now no one can, and unlike almost everything else in this story, that part is entirely within your power to do better.
Every argument in this piece lands on the same missing artefact: a record of the decision that an outsider can inspect. Chain of Consciousness is that record for agent work — a verifiable log of what an agent saw, what it decided, and on what basis, so a judgment call travels with its own evidence instead of being asserted afterwards.
pip install chain-of-consciousness · npm install chain-of-consciousness