"Agent" answers the wrong question

Two systems ship this quarter, and both teams call them agents. The first drafts replies to support tickets; a person reads every word before anything leaves the building. The second closes stale accounts, issues refunds under a threshold, and emails the customer about it, with nobody watching in real time. Same word, same diagram with arrows looping through a model, same slide at the all-hands. Everything that matters about these two systems is different, and the shared vocabulary cannot say so.

"Agent" describes ambition. It tells you what the builder hopes the system will feel like. It does not answer any question an engineer, an operator, or a buyer actually needs answered: which commitments may this software make without a person, which of those can be undone, and who answers when one of them is wrong. The industry's reflex has been to tighten the definition — a model in a loop with tools, autonomous pursuit of a goal, software that plans. These definition fights argue about the boundary of a category, and the useful information was never the category. It is a coordinate inside it: the specific decisions this specific system is permitted to commit on its own. Whether something is an agent is a marketing question. What it may decide is an engineering one.

The unit of delegation is a decision, not a task

Delegation goes wrong at the moment of description, because systems get described by task. Handle inbound support. Manage renewals. Do the bookkeeping. A task is a bundle, and inside every bundle sit decisions of wildly different weight. Handle inbound support unpacks into: decide what the customer is actually asking, decide whether their account qualifies for a remedy, decide which remedy, decide the wording, decide whether a human needs to see this one. Grant autonomy at the task level and you have granted authority over every decision in the bundle, including the ones nobody said out loud.

This is where most disagreement about AI features actually lives. When one person says the feature is ready and another says it is reckless, they are rarely disagreeing about the model. They are holding different unwritten inventories — one pictures the software deciding the wording, the other pictures it deciding the remedy — and because the conversation happens at the task level, neither inventory ever reaches the table.

We made a version of this argument about analytics: a verdict is a smaller promise than a dashboard turns on the moment software stops presenting information and starts committing to a claim. Delegation is the same line, drawn inside an AI-native system, decision by decision: for each item in the inventory, is the model informing a commitment a person makes, or making the commitment itself? Write that inventory before the prompt. The prompt is downstream of the authority grant, and no amount of prompt engineering repairs an authority grant nobody wrote down.

Four words that do the work "agent" refuses to

The vocabulary that works is small, and it attaches to decisions, not to systems.

Suggests. The model's output is an input to a person's decision, and the person commits. The obligation is legible reasoning and cheap dismissal. A suggestion that is expensive to ignore — buried defaults, nagging re-prompts, extra friction on "no" — is a decision wearing a costume.

Drafts. The model produces the artifact; a person edits it and signs it. The person still commits, but the default has moved, and defaults are where delegation actually happens. The obligation is a review surface on which editing is easier than approving, because a reviewer who only ever approves has been quietly promoted out of the decision while keeping the title.

Acts, with undo. The model commits, and people audit afterward. The obligation is an undo that genuinely restores the prior state rather than an undo-shaped button, an audit trail written for a human reader, and a review cadence that does not depend on somebody remembering to look.

Acts, for keeps. The model commits things that cannot be taken back — the money moves, the message sends, the record is gone. The obligation is hard bounds stated in the domain's own terms, escalation triggers that fire before the boundary rather than after it, and a named person who answers for the outcomes. Not a team. A name.

The point of the ladder is that one system holds different rungs for different decisions. The same product might act for keeps on scheduling, draft the outreach, and only ever suggest on pricing. Ask whether that product is an agent and you have flattened a map into a syllable.

Every rung is a bill, not a badge

Moving a decision up the ladder is routinely announced as a feature. It is a purchase. Suggests to drafts buys the review surface. Drafts to acts buys the undo, the audit trail, the sampling routine. Acts-with-undo to acts-for-keeps buys the bounds, the escalation, and the name. Teams that adopt "agent" as an identity rather than a coordinate tend to build the autonomy first and receive the invoice in production, where it arrives itemized as incidents — the missing undo discovered by the first commitment worth undoing.

The ladder is also not a maturity model, and this is the misreading the word "agent" most encourages. The right rung for a decision is not the top one; it is the rung where the value of not waiting for a person outweighs the cost of a wrong commitment, discounted by how completely that commitment can be reversed. Plenty of decisions should live at "suggests" indefinitely, and naming rungs is what makes staying put an argument instead of an apology. "This stays at drafts because a wrong commitment costs a customer relationship" is a position an engineer can defend in review. "It's not really agentic yet" is a confession of roadmap guilt about a system that may already be exactly right.

Scope written in these words is a different negotiation

A request for "an autonomous agent" is almost always a request for a rung that has not been priced, by the person asking or by the team agreeing. The fit criteria published at the studio ask for room to challenge scope before committing to output, and on an AI-native build this is the first challenge worth making: not which model, not which framework, but which decisions, at which rung, answered for by whom. Our refusal list includes delivery plans built around an immovable feature list, and "make it fully autonomous" is a feature list with one immovable item. The honest counter-move is the decision inventory, rung by rung, with the bill attached to each.

There is a pattern here the studio has committed to elsewhere. The ventures catalogue stays empty until products are ready to show; no email address is published until the mailbox is live; we publish the work we decline so that the work we accept means something. All of these are the same move: writing down the boundary of the promise so the promise can be checked. "Acts with undo on scheduling, drafts on outreach, suggests on pricing" is a far smaller promise than "autonomous agent," and it is the one that can be kept, audited, and grown. Our working loop ends in Compound, and compounding requires exactly this: a record of which commitments the system made and how they turned out. A system whose authority was never enumerated cannot produce that record, because nobody can say which of its actions were its to take.

The question that opens an AI-native build is not "should this be an agent." It is: here is the list of decisions inside the task — say, for each one, who commits. Every answer to that question is buildable. The word "agent" was never an answer to it.