The phrase answers where and skips everything else
"Human-in-the-loop" is a statement about topology. It says that somewhere between a model producing an output and that output taking effect in the world, a person appears. That is all it says. It does not say what the person sees, what they are allowed to do, how long they have, or who answers when the thing they waved through turns out to be wrong.
Those omissions are the whole design. A review that shows the reviewer the wrong evidence is theatre. A review whose only action is "approve" is a signature. A review whose time budget is set by the queue behind it converges on whatever keeps the queue moving. A review whose reviewer carries the blame for the model's mistakes is a liability transfer dressed as a safeguard.
The phrase survives because it sounds like a safeguard. It arrives in the design conversation and ends it. Someone asks what happens when the model is wrong, someone says there is a human in the loop, and the conversation moves on. This is the same move that "agent" makes, and the cure is the one argued in stop calling it an agent and say what it may decide: stop naming the shape and start naming the decision.
What the reviewer actually sees
Start with the input, because most loops fail here before anyone clicks anything.
The model saw evidence: a document, a transaction history, a conversation, a set of readings. It produced an output: a label, a draft, a recommendation, an action. The question is whether the human sees the evidence or the output. Almost always it is the output, sometimes with a short rationale the model also produced.
A reviewer who sees only the output is not checking the model. They are doing the same job as the model, with less information. They are pattern-matching on the surface of a result, judging whether it looks like the kind of thing that is usually right. The model already did that, and it did it with the evidence in front of it. Putting two pattern-matchers in series, the second one blindfolded, does not create a check. It creates a second opinion about plausibility.
The rationale makes this worse, not better. A generated explanation is another output, produced by the same process. It reads as justification and functions as persuasion. A reviewer handed a confident, well-formed explanation has been given a reason to approve, not a means to verify. This has a sibling in hallucination is a word that protects the builder: the explanation is not the model lying, it is the model doing exactly what it was asked, and the design is what put the reviewer in a position to be persuaded by it.
So the first honest sentence to replace the phrase is: the reviewer sees this evidence, in this form, alongside the model's output. If that sentence cannot be written, there is no review, only a second guess.
What the reviewer is allowed to do about it
The next omission is the action set. Draw the interface. Count the buttons.
If the buttons are approve and reject, ask what reject does. If rejection routes the item back to the same model with the same evidence, the reviewer has a button that changes nothing. If rejection drops the item into a queue nobody owns, the reviewer learns quickly that rejection makes work disappear rather than get fixed, and they stop rejecting. If rejection requires a written justification and approval requires a click, the design has priced the two actions differently, and the price is what people respond to.
If the reviewer can edit the output, the loop has changed character. The model is now drafting and the human is authoring, and every edit is a signal that should reach whoever tunes the system. If those edits go nowhere, the design has hired an editor and thrown away the edits.
If the reviewer can escalate, ask to whom, and whether that person has a different action set. Escalation to someone with the same buttons is a delay with a title.
The action set is where the design either grants authority or withholds it. A person who cannot change the outcome is not in the loop. They are beside it, and the diagram is lying about the arrow.
The rubber stamp is what the design asked for
When a loop degrades into approval-by-default, the tempting explanation is that the reviewer got lazy. This is wrong, and wrong in a way that matters, because it blames the person for the structure.
Consider what the reviewer experiences. Most outputs they see are fine, because the model is mostly fine. Their attention calibrates to that base rate. The interface presents approval as the default path, because that is the path most items take. The queue does not shrink on its own, and the queue is visible. Above them, someone is watching how long items sit in review, because review time is a delay the product carries.
Every one of those pressures points the same way. Approval gets faster. Rejection gets rarer. Eventually the approval rate settles at a value nobody chose, and that value becomes the system's effective accuracy. This is the mechanism described in every threshold is a decision someone stopped making, operating on a person instead of a parameter. The reviewer's tolerance is the threshold, and it drifted because no one wrote down what it was supposed to be.
The consequence is that a human in the loop is not a fixed property of a system. It is a rate, and rates decay under load. Any design that depends on the loop has to say what keeps it from decaying: a sampled second review, a cap on items per reviewer, a rule that some of the approved items are re-examined with the evidence visible. Without a mechanism, the loop is real on launch day and ornamental within the quarter, and no one will be able to say which day it crossed over.
Shapes that can be stated honestly
If the phrase has to go, it should be replaced with a shape that has a stated cost. There are only a few.
The gate. The person decides before the effect. They see the evidence, not just the output. They can reject, and rejection goes somewhere. Throughput is bounded by what a person can genuinely examine, and that bound is the point of the design. If the business case requires more throughput than the gate allows, the business case requires a different shape, and saying so early is the honest move.
The audit. The model acts. A person examines a sample afterward, with the evidence, and with the power to stop the system. This is a loop only if the findings flow back: into the model's scope, into the thresholds, into the cases that get routed elsewhere. An audit whose findings are filed is a report, and a report is not a loop. The person in this shape needs standing to halt the system, and that standing needs to be written down before the first incident, not negotiated during it.
The envelope. The model acts inside a defined boundary and defers everything outside it. The person sees only what was deferred. This is the shape that "agent" should have meant all along: a named list of what the system may decide, and a named person for the rest. The work is in drawing the boundary, and the boundary has to be drawn in terms of consequences, not confidence. A model that defers when it is unsure has been given a boundary it controls. A model that defers when the outcome is irreversible has been given one that someone else controls.
Each of these shapes costs something specific: capacity, latency, or scoping effort. "Human-in-the-loop" is attractive precisely because it lets a design claim the benefit of all of them without paying for any.
Deciding who sits there is a scope decision, not a UX one
None of this can be settled by the people building the system. Who reviews, what they may veto, and what they are accountable for are questions about an organisation, and the answers belong to whoever owns the outcome.
This is why the fit criteria on the studio ask for direct access to decision-makers and room to challenge scope before committing to output. A loop designed without the owner in the room will be designed around the owner's absence. It will assume a reviewer exists, assume they have time, assume they have authority, and ship. Operation will reveal which assumptions were wrong, by which time they are in the architecture. The same reasoning runs through the handoff is an architecture decision: the people who will operate the loop are a design input, not a deployment detail.
It is also why the refusal list on the same page includes work that depends on fabricated proof. A person placed in a loop to satisfy a checklist, with no evidence in view and no button that changes anything, is fabricated proof of oversight. It exists so that the diagram has a human on it. When the system fails, the diagram will be used to show that a human was there, and the human will be the one explaining why they approved it. Declining to build that is not caution. It is refusing to build a mechanism whose function is to assign blame after the fact.
An operating thesis that keeps operating what it launches makes this concrete. Whoever runs the system inherits the loop, the queue, the reviewer's drift, and the incident that eventually arrives. That operator has every reason to want the loop real and the vocabulary exact. So the phrase should not appear in a design document. In its place: what the reviewer sees, what they may do, what keeps their judgement from decaying, and who answers when it fails. A handful of sentences, each of which can be checked. That is a smaller promise than the phrase, and as argued in a verdict is a smaller promise than a dashboard, the smaller promise is the one that gets kept.