← Back to blog

Embeddings Are Saussurean, Which Is Exactly Why They Hallucinate Reference

A well-formed citation that refers to nothing is not a mistake in the ordinary sense. It is a system doing exactly what its theory of meaning tells it to do, a theory written down in 1916 with no room for the world, on purpose.

Published July 2026 · 12 min read · embeddings / semiotics / hallucination / symbol grounding


In the spring of 2023, a New York lawyer named Steven Schwartz turned to ChatGPT for his legal research, and the brief that resulted cited a case called Varghese v. China Southern Airlines. It was a good citation. It had a plausible name, a plausible court, a docket number and a year, and it quoted at length from the opinion. It cited other cases in turn, in the correct format, with the correct cadence of legal prose. There was only one problem, and it was total: the case did not exist. Neither did the five other cases the brief leaned on. When opposing counsel could not find them and Schwartz asked ChatGPT whether they were real, the model assured him they were. That June, a federal judge named P. Kevin Castel sanctioned Schwartz and his firm five thousand dollars, and the episode became the textbook example of what we have learned to call an AI hallucination.

Here is the thing worth staring at, though, before you file it under "the model made a mistake." Varghese v. China Southern Airlines is not a mistake in the ordinary sense. It is a perfectly well-formed citation that refers to nothing. Every relation it has to other legal signs is correct. It looks like a case, sits among cases, quotes like a case, cites like a case. The only thing it lacks is the one thing a citation is for, which is a real case at the other end of it. And a system that can produce that fluent, coherent, referent-free sign is not malfunctioning. It is doing exactly what its theory of meaning tells it to do, a theory written down in 1916 by a Swiss linguist who never saw a computer, and who built, on purpose, a model of meaning with no room for the world.

The 1916 model with no slot for the world

Ferdinand de Saussure died in 1913, and three years later two of his colleagues assembled his lectures into the Course in General Linguistics, a book that founded modern linguistics from beyond the grave. Its central move is still one of the most radical ideas in the human sciences, easy to state and hard to fully believe: a sign's meaning is not in what it points to. It is in how it differs from every other sign.

Saussure's own line, in the standard translation, is stark. "In language there are only differences without positive terms." A word has no positive content of its own. Its value is entirely a matter of its place in the system, of what it is not. "Cat" means what it means not because of anything intrinsic to those three letters, but because it is not "cot," not "dog," not "kitten." The meaning is the difference, and nothing but the difference.

Two features of this model carry the weight for everything that follows. The first is that Saussure's sign is a dyad, a pairing of exactly two things: a signifier, the sound or the spelling, and a signified. And the signified, crucially, is a concept, a mental idea, and not the thing in the world. The word "tree" binds a sound-pattern to a notion of tree-ness in your head. Notice what is missing from that pair: the actual tree, the real object out in the yard. Saussure left it out deliberately. He was adamant that "language is not a nomenclature," not a list of names hung on a collection of pre-existing things. The naive picture, where words are labels and the world supplies the objects to be labeled, is precisely the picture he built his system to replace.

The second feature follows from the first. Because meaning is difference within a closed system, the system never has to leave itself. Every sign is defined by other signs. The whole apparatus turns inward, sign against sign, and the referent, the actual thing in the world, sits outside the model entirely, by design. As a science of language treated as a self-contained system, this was a brilliant and productive abstraction. It is also, if you build a machine on it, a hole shaped exactly like reality.

An embedding is that model, rebuilt in floating point

Now the part that should make a linguist and a machine-learning engineer glance at each other. A vector embedding, the representation at the heart of every modern language model, is Saussure's theory rebuilt in linear algebra, feature for feature.

An embedding represents a word as a position in a high-dimensional space, and that position is learned entirely from the company the word keeps. The idea has a name in linguistics, the distributional hypothesis, and a slogan from the linguist J.R. Firth in 1957: "you shall know a word by the company it keeps." Zellig Harris had made the same case, more formally, in 1954. When word2vec (Mikolov and colleagues, 2013) and GloVe (Pennington and colleagues, 2014) turned that slogan into an algorithm, what they built was Saussure operationalized: a word's meaning is its pattern of differences from other words, now written as coordinates. The celebrated party trick of those early embeddings, that the vector for "king" minus "man" plus "woman" lands you near "queen," is neither magic nor understanding. It is Saussurean value made arithmetic. The words have no intrinsic content; they have only relations, and the relations turned out to be regular enough to add and subtract.

Which means an embedding has exactly what Saussure's sign has, differential position in a system, and it is missing exactly what Saussure's sign is missing, a slot for the referent. Had the model built a vector for "Varghese v. China Southern Airlines," that vector would be defined by its neighbors, other case names, aviation-law phrasings, citation formats, and not one of those neighbors is the actual case, because there is no actual case, and nothing in the geometry could tell the difference. The embedding does not know that the case is fake, in the same way and for the same reason that it does not know any word is real. Reality was never one of its coordinates.

Hallucination is the machine working exactly as built

Follow that through and the whole phenomenon of reference hallucination changes character. A purely relational system, by construction, optimizes for one thing: relational coherence, the sign whose value fits the system. It produces the distributionally plausible continuation, the citation that sits right among real citations, the case name shaped like a case name, the quotation shaped like a quotation. Whether that sign connects to anything real is not a question the system gets wrong. It is a question the system cannot ask, because its founding model has no term for the real, no operation that compares a sign to the world, no place where such a comparison would even live.

So reference hallucination is not a defect layered on top of a working system. It is the system working, coherently and as designed, in the one regime where coherence and truth come apart. The astonishing thing, once you see this clearly, is not that a language model sometimes invents a case that does not exist. It is that a purely relational system, a machine that has only ever seen words in the company of other words and has never once touched the thing a word is for, manages to refer correctly as often as it does. And when it does refer correctly, that is not the system reaching out to the world. It is the distributional pattern happening to line up with the world, a correlation rather than a connection. The correlation is often excellent, because the training text was written by people who were themselves connected to the world. But it is borrowed reference, secondhand, and it fails exactly where you would predict: at the specific and the verifiable, the case name and the docket number and the dollar figure, the places where a sign has to be pinned to one particular thing and there is no pin.

The field rediscovered the hole and renamed it

Artificial intelligence did not inherit this blind spot knowingly. It rediscovered it, from the inside, and gave it a new name. In 1990, the cognitive scientist Stevan Harnad published a paper in Physica D called "The Symbol Grounding Problem," posing a question that is Saussure's model turned into a trap. A purely symbolic system defines each of its symbols in terms of other symbols. How, Harnad asked, does any symbol ever connect to what it means? His image was of a person trying to learn Chinese from a Chinese-to-Chinese dictionary with no prior Chinese: you look up a symbol, its definition is more symbols, you look those up, their definitions are more symbols, and you circle the dictionary forever without ever breaking out to the world. Meaning defined purely by other meaning is a closed loop with no exit.

That closed loop is Saussure's dyad exactly, described seventy-four years later by a different field that had walked into the same wall. Signs defined by signs. Meaning by difference. No door to the referent. The symbol grounding problem of 1990 and the relational model of 1916 are not two problems; they are one gap, seen twice. And that is the reframe that matters for anyone building with these systems. Reference hallucination and the grounding problem are not novel bugs of neural networks, to be patched in the next release. They are the structural inheritance of a theory of meaning older than the transistor, and they will not be tuned away, because they are not a matter of the relations being slightly off. They are a matter of a missing term.

Grounding is not a tune-up, it is a third leg

This is why every serious fix for hallucination has the same shape, and why none of them is a cleverer prompt or a larger model. Retrieval-augmented generation, tool calls, sensors, multimodal perception: every one bolts the relational system to something outside it. Retrieval attaches the model to a real corpus it can be checked against. A tool call attaches it to a real system that answers with real state. A sensor attaches it to a real environment. In each case you are not improving the relations inside the embedding; you are adding a term the embedding never had. Harnad himself pointed this way in 1990, proposing to ground symbols from the bottom up in non-symbolic, sensory experience. You cannot train reference into a relational space, because reference is not a relation the space contains. You have to connect the space to the world and supply the missing thing.

Semiotics drew that missing thing explicitly, and the diagram is worth seeing. We have written before about Charles Peirce, Saussure's great contemporary, whose sign is not a dyad but a triad: sign, object, and interpretant. Peirce kept the leg Saussure dropped. His "index" is a sign that is, in his phrase, "really affected by" its object, welded to the world by causation, the smoke that cannot lie about fire. And in 1923, Ogden and Richards drew what is now called the triangle of reference: a symbol at one corner, a thought at another, a referent at the third. The two sides that connect symbol to thought and thought to referent they drew as solid lines. The bottom side, the one directly joining the symbol to the referent, they drew as a dotted line, because that connection is not given. It is indirect, it runs through the world, and it is the line most easily broken. A raw embedding is Saussure's dyad: two corners and no referent, no bottom line at all. An embedding grounded by retrieval is that dyad handed its third corner, the dotted line finally drawn in, from outside.

The honesty here matters. Grounding does not settle reference the way a proof settles a theorem. Retrieval can fetch the wrong document, a tool can return stale state, a sensor can be miscalibrated. Bolting on a referent introduces a whole new surface of error. But it is the right surface, the one place where a sign meets a thing and can be checked against it, and that surface does not exist inside the embedding at all. The point is not that grounding is perfect. It is that grounding is the only move with the right shape, because it is the only move that adds the missing term instead of rearranging the ones you already have.

What to do with a machine that has never met the world

So here is the practical residue for anyone who builds with embeddings, and it is mostly a change in what you expect from them.

Stop treating reference hallucination as a tuning problem. You will not prompt it away, and you will not fine-tune it away, because it is not an error in the relations. It is the absence of a referent, and no adjustment to a system of differences will ever add one. A model that invents a citation does not need more training on citations. It is a relational system doing relational work in a spot that demanded a connection to the world.

Learn to tell your relational operations from your referential ones, and never let the first impersonate the second. An embedding similarity, a nearest-neighbor lookup, a generated continuation: these are all relational. They tell you what fits, what is plausible, what belongs in the neighborhood. They do not tell you what is true, and they cannot, because truth is a relation between a sign and a thing, and there is no thing in the box. The instant your question becomes "is this real," you have left the territory the embedding can speak about, and you need a referent, which means going out to touch the world: a lookup against a real record, a call against a real system, a measurement of a real state.

And treat grounding as architecture, not garnish. The referent is not a feature you sprinkle on a finished model to shave a few points off its error rate. It is the leg the whole structure is standing without, and where and how you attach it decides what your system is honestly able to be about. Build the connection to the world in on purpose, at the center, or the model will go on producing signs that are beautifully coherent and quietly about nothing.

Saussure was not wrong. His relational account of meaning is one of the deepest things anyone has ever noticed about language, and embeddings are its vindication in a domain he could not have imagined. But he built a theory of meaning with no place for the world, deliberately, as an abstraction, and when we rebuilt that theory in linear algebra we inherited the hole along with the genius. An embedding, exactly as Firth promised, knows a word by the company it keeps. What it has never done, not once, is meet the thing the word is for. If you want it to be about the world, that introduction is your job, and it always will be.


Sources

An embedding tells you what fits. It cannot tell you what is real. The instant your question becomes "is this true," you have to go touch the world.

A relational model optimizes for the sign whose value fits the system, and its confident output is exactly as fluent when it refers to nothing. The only move with the right shape is to add the missing term: check the sign against a real referent. That is what the agent trust stack is built to do, checking an agent's output against ground truth rather than trusting the coherent shape, keeping a provenance record of what the agent actually did against real state, and pricing an agent by a track record measured against reality rather than against its own plausibility. Grounding is architecture, not garnish; build the connection to the world in on purpose.

Read the Theory of Agent Trust

pip install agent-trust-stack  ·  npm install agent-trust-stack

Or the provenance record on its own: pip install chain-of-consciousness / npm install chain-of-consciousness.