← Back to blog

The Design of Everyday Agents

A Norman door tells you the truth the moment you push it. An agent's wrong answer arrives in the same prose as a right one, at the same speed, under the same green checkmark.

Published September 2026 · 12 min read · interaction design / agents / evaluation / Don Norman


In February 2016 the podcast 99% Invisible and Vox released a short video whose title has since become a small genre of its own: "It's not you. Bad doors are everywhere." The subject was the door that carries Don Norman's name, the one with a handle that asks to be pulled and then wants a push, or a flat plate on the side that opens the other way. Norman had been complaining about it since 1988, when he published The Psychology of Everyday Things, later retitled The Design of Everyday Things and revised in 2013. The Norman door is the book's villain. It is the object people reach for when they need an example of a design that fails the person using it.

I want to defend the door. Not because it is good; it is not. Because on Norman's own criteria it is a better-designed object than the chat box you probably used this morning, and seeing why tells you something about agents that the usual vocabulary does not.

Push a Norman door the wrong way and it does not open. That is the whole failure, and look at its properties. It is instant. It is unambiguous. It is honest: the door told you the truth about the state of the world the moment you acted on it. You are confused for two seconds, you form an accurate model, "pull," and you are through. The door's failure is legible at the moment of failure, and the entire cost of the failure is the two seconds.

Now ask an agent for something and get a wrong answer. It arrives in the same fluent, well-structured prose as a right one, at the same speed, with the same confidence, under the same green checkmark. There is no equivalent of the door not opening. You may discover the failure an hour later, or a week later, or never. The Norman door has an almost zero gulf of evaluation. The chat box has one you could lose a project in.

That phrase, gulf of evaluation, is Norman's, and it comes from the half of his book that nobody borrowed.

The half of the book nobody took

The words everyone took from Norman are affordance and, after the 2013 revision, signifier. He added the second because the first had been misused; in ACM Interactions in November 2008 he wrote that "the term has been widely used and misused," and that what designers needed was perceptible clues rather than a theory of possibility. Those two words are now furniture. We have used them ourselves, more than once.

The half nobody took is Chapter 2 of the revised edition, "The Psychology of Everyday Actions," with its two gulfs and its seven stages.

The Gulf of Execution is the distance between what you want and the actions the thing offers you to get it. His 1988 example was a film projector that needed a long, unobvious threading sequence to do the one thing anyone wanted from it. The Gulf of Evaluation is the effort it takes to work out, from what the system shows you, whether what you wanted actually happened. The same projector gave no way to tell whether the film was threaded right until you ran it.

The seven stages, in the 2013 edition's one-word labels: goal, plan, specify, perform, perceive, interpret, compare. The 1988 wording is longer, and the research literature still uses it: establishing the goal, forming the intention, specifying the action sequence, executing the action, perceiving the state of the world, interpreting it, and evaluating the outcome against the goal. Norman's own division is the part that matters here. Plan, specify and perform bridge the Gulf of Execution. Perceive, interpret and compare bridge the Gulf of Evaluation. Goal sits above both, at the top of the loop.

Credit where it is due, and early. "Norman's gulfs describe LLM interfaces" is not a new observation; it is published research. In September 2023, revised in March 2024, Hariharan Subramonyam, Roy Pea, Christopher Pondoc, Maneesh Agrawala and Colleen Seifert posted "Bridging the Gulf of Envisioning," whose abstract says that our limited grasp of how a person goes from a goal to an intention is "a blindspot even in established interaction models such as Norman's gulfs of execution and evaluation." Their answer is a third gulf, the Gulf of Envisioning, sitting between Norman's first stage and his second, with three gaps inside it: a capability gap, "the users' inability to formulate 'how to' procedures to implement their intentions"; an instruction gap, "the user's challenges in clearly and effectively expressing their intentions in the interface as text prompts"; and an intentionality gap. And in defining that third gap they noticed the thing this essay is about. In the simplest path through a language model, they write, "the user states their goal directly and forgoes any cognitive task processes anticipated in doing the tasks themselves (i.e., planning what to write, executing it, and evaluating it). Instead, they start with evaluation of the LLM-generated text."

What the paper does with that observation is add a stage in front of Norman's model, and that is a fair thing to do with it. What I want to do is follow Norman's existing stages after they move. They do move, and they move in a very particular way.

Where the stages went

Norman's model assumes that one actor performs all seven stages. You form the goal, you plan, you specify, you act, you perceive, you interpret, you compare. The seven stages are one person's loop, which is why the book can talk about "the user" in the singular and mean the whole cycle.

An agent splits the loop between two parties. Assign each stage to whoever actually does it:

Stage2013 labelWith an agent, done by
1goalthe human
2planthe agent
3specifythe agent
4performthe agent
5perceivethe human
6interpretthe human
7comparethe human

Look at where the cut falls. Stages two through four are, in Norman's own division, the entire Gulf of Execution. Stages five through seven are the entire Gulf of Evaluation. The line between human and agent runs exactly along the seam of the first gulf. An agent does not bridge the Gulf of Execution. It annexes it. The three stages that constitute that gulf move to the machine wholesale, and the three stages that constitute the other gulf stay where they always were.

This is a reading of Norman's model, not a measurement. No study has weighed it, and I am not going to pretend one has. But it explains something the products are sold on and the users keep walking into. Annexing execution looks like a pure win, and in one sense it is; the agent plans, specifies and performs faster than you, and often better. The trouble is on the far side of the seam. You are now perceiving, interpreting and comparing the outcome of a plan you did not make, a specification you did not write, and actions you did not watch. Each of those is information that Norman's model assumes you hold when you arrive at stage five, because in his loop you have just done stages two through four yourself. With an agent, you have not. So the Gulf of Evaluation did not keep its width when execution left. It got wider because execution left, and it keeps getting wider as the agent improves, because the better the agent is at the three stages it took, the less inclined you are to look hard at the three you kept.

Put the door beside it once more. The door leaves you all seven stages and makes stage five trivial: it did not open. The agent takes three stages off you and makes stage five the hardest thing in the loop.

Norman's remedies, turned on the subject

Norman did not only name the gulfs. He said what bridges the evaluation gulf, and it is two things: feedback, and a good conceptual model. In the fundamental principles that open the 2013 edition, feedback has to be immediate and it has to be informative, and a conceptual model is the simplified account of how the thing works that the design helps you build in your head. Agents are weak at both.

Feedback first. A spinner and then a wall of prose is feedback about completion. It tells you the system finished. It tells you nothing about whether the system was right, and those are two different signals presented identically. The Norman door's feedback is about correctness; it reports the state of the world. The agent's feedback reports the state of the agent. Norman asked for informative, and "done" is the least informative true thing a system can say.

Conceptual model second, and here the conversational surface does something worse than fail. It installs a model, and the model is wrong. A thing that answers in the first person, in complete sentences, with reasons, invites you to model it as a colleague who understood the task. That is the one model guaranteed to make your stage-six interpretation fail worst, because you will read confidence as competence.

We do not have to guess what Norman thinks of this. He has said it, about this technology, on the record. In a Design Observer conversation recorded at the Design Research Society's 2024 conference in Boston, in June 2024, he said: "the large language models have zero understanding." And: "They're powerful ... but they do pattern matching and they don't understand what they're putting together." He offered a conceptual model of his own in the same conversation. Working with one, he said, will be "like you just hired as an assistant who was really, well, naive," and then, "maybe the assistant will make things up too, I don't know." Two months earlier, talking to Doc Searls, he had compressed it to an aphorism: "Artificial is AI's first name. And Intelligence is a quality, not a quantity."

Set his model against the interface's. The naive new assistant who may make things up is a model that tells you to check. The fluent colleague is a model that tells you not to bother. Both are conceptual models in Norman's sense. The first one narrows the gulf and the second one does not.

The clue that points the wrong way

Norman's 2008 essay opens with "We are all detectives, searching for clues to enable us to function in this complex world," and defines a signifier as "some sort of indicator, some signal in the physical or social world that can be interpreted meaningfully." I will use the word sparingly, since we have spent it elsewhere, but it earns a paragraph here.

A chat box offers one signifier for input, a blinking cursor, which says "type anything." It is the least informative clue available, attached to the most capable and most variable system most people have ever used. The flat plate on a door at least narrows the world to one action. And on the output side the clue is worse than absent. The surface features a reader uses to judge whether a text is reliable, its confidence, its structure, its specificity, its register, are the features a language model produces whether or not the content is right. The clue points, and it points the wrong way. To be exact about Norman's terms, because he added the second one in 2013 to stop people blurring it with the first: agents have no shortage of affordances. They can do an enormous number of things. What they lack is signifiers for which of those things just happened and whether it worked.

The evidence, and the old diagnosis

None of this is only theory. In September 2025, revised in May 2026 and accepted at the 2026 ACM Conference on AI and Agentic Systems, Pradyumna Shome, Sashreek Krishnan and Sauvik Das published "Why Johnny Can't Use Agents," the title a nod to the 1999 usable-security paper about PGP. They reviewed the marketed use cases of 102 commercial AI agents and found they sort into three categories: orchestration, creation and insight. Then they put 31 participants in front of two of the best-known agents, Operator and Manus, on representative tasks from each category. The participants, in the paper's words, "were generally impressed with these agents but faced significant usability challenges ranging from agent capabilities that were misaligned with user mental models to agents lacking the meta-cognitive abilities necessary for effective collaboration."

Read that against the stage table. A misaligned mental model is a stage-six failure: interpretation run on the wrong conceptual model. Missing meta-cognition is the agent's inability to report on its own stages two through four, which is the very information the human needs at stage five and no longer has. The two named obstacles are the two sides of the seam.

Norman's founding move, in 1988, was to insist that when a person struggles with an object, the fault is in the design and not in the person. That is why the door got his name and not the name of anyone who pushed it. The diagnosis is available again and it is mostly unused. When a user misjudges what an agent did, the industry's instinct is to talk about prompting skill, which is a way of saying the person pushed wrong. Norman would look at the door.

So here is the practical residue, for anyone building one of these. You cannot give the human back stages two through four; taking them is the product. What you can do is stop treating stage five as solved. Report correctness separately from completion, and make the two look different on the screen. Have the agent narrate the three stages it took, the plan, the specification, the actions, in a form short enough to be read and specific enough to be checked, because that narration is the only evidence stage six now runs on. Install Norman's model instead of the interface's: a capable, naive assistant who may have made something up, which is a model that tells the person to look. And judge the result by his oldest test, the one the door passes with a flat plate and no moving parts, and the chat box fails: when it goes wrong, does the person know at the moment it goes wrong? That is a low bar. Most agents, today, do not clear it.


Sourcing notes: Norman's remarks on language models are quoted from the Design Observer "Design As" conversation recorded in June 2024 and from Doc Searls' 5 April 2024 post, and are attributed to those dates and not to the book; the seven-stage labels are the 2013 revised edition's, with the 1988 wording given as the Envisioning paper renders it in its Section 2; the reallocation of stages between human and agent is an analytical reading of Norman's model and no study is claimed to have measured it; the Envisioning paper's three gaps and its "forgoes ... Instead, they start with evaluation" passage are quoted from its Sections 3.2.1 to 3.2.3, and the essay states plainly that the paper made that observation first; the Johnny paper's figures (102 agents, 31 participants, Operator and Manus) are from its abstract; no page numbers from the book are cited, and no direct quotation from the book is made.

Sources

The narration is the only evidence stage six now runs on.

If the practical residue above is right, an agent has to report the plan, the specification and the actions it took in a form short enough to read and specific enough to check. Chain of Consciousness is that record made tamper-evident: written at the moment of the decision rather than reconstructed afterward, so the person at stage five is reading what happened instead of what the agent says happened.

Hosted Chain of Consciousness

pip install chain-of-consciousness  ·  npm install chain-of-consciousness

Or the whole stack, provenance and ratings and verification together: pip install agent-trust-stack / npm install agent-trust-stack.