Writing Your Agent's Commercial Authorization Scope
If someone asked what your agent is authorized to buy, could you hand them a document? Here is what belongs in a commercial authorization scope, and the two measurements that decide its shape: how concentrated agent payment flow actually is, and how many endpoints have ever had a second buyer.
· 989 words
If someone asked today what your agent is authorized to buy, could you hand them a document? Not a description of what it does in practice. A record, written before it launched, of what you decided it was permitted to do.
The Federal Trade Commission has opened a consumer protection investigation into OpenAI, Anthropic, and other large AI labs, with Civil Investigative Demands expected to go out within weeks. In late September the agency had already staked out the position that a developer answers for what its agents do. Whatever becomes of that particular investigation, the direction it points is clear enough to design against: the party that deployed the agent is the party asked to explain it. A written scope is what you explain with.
Here is what belongs in one, and the two measurements that decide its shape.
Payment boundaries: name the head, write policy for the tail
The first instinct is to enumerate. Name every service your agent may pay, review each one, done. Whether that works depends on how the spending is actually distributed, which is measurable.
Our index records 1,311 distinct services receiving USDC payment in the trailing week. Taken alone, that count argues against enumeration. But the flow is heavily concentrated rather than broad-based: the single largest recipient accounts for 34.5% of all measured transfers. One service out of 1,311 carries better than a third of the volume. This is a rolling one-week window, not a cumulative total, and it should not be described as broad-based activity.
That concentration is the useful fact, and it points the opposite way from the headline count. A named allowlist holding a small number of entries covers a disproportionate share of what your agent actually spends, at very little review cost. Enumeration works fine for the head of this distribution.
So the payment boundary section has two halves, and the second is where the thinking goes. First, name the services your agent is pre-authorized to pay. That list will be short. Then write the policy governing everything else: the per-call price ceiling that needs no further approval, the cumulative daily ceiling, the value above which a person clears a payment before it fires, and which classes of service the agent may reach into without a name on the list. The tail is where an unreviewed counterparty enters your system, so the tail is what the policy exists for.
Write those thresholds down before you implement them. The code enforces whatever you built. The document is what establishes whether what you built was what you intended.
The pre-payment check, and why 18.2% sets its bar
Your scope should state what checking happens before your agent pays a service for the first time. How thorough that needs to be depends on whether anyone else has already vetted the service, and that is also measurable.
Of 40,466 listed paid endpoints, 7,382 have more than one distinct paying buyer, which is 18.2%. That measurement comes from Coinbase telemetry we ingest rather than from our own probe, so read it as their view of their own network.
Now read it from the other direction. For a listed endpoint outside that 18.2%, no second external party has committed money to the service. There is no other buyer whose diligence you could lean on, and no repeat custom from anyone except the operator. That is the condition your pre-payment check has to be written for, rather than the 18.2% that already has a second buyer.
What the check contains is your call, but write it specific enough to actually run: a liveness handshake, a scope declaration fetched and compared against what your agent intends to call, a price confirmed before the first payment settles, and a record of who approved adding that service or its category.
The action inventory
Payments are one class of consequence. List the others. Every action your agent can take that changes something outside its own context belongs here: sending a message, placing an order, booking a slot, submitting a form, paying an invoice, starting a subscription.
Against each one, record what triggers it, whether it can fire without a person, what the person sees before approving where approval is required, and what happens when approval is refused. That last one is the entry teams skip. A refusal is a legitimate end state for that branch of work, not a failure to retry around.
This inventory is your statement of intent. When an agent does something a user did not want, the first question is whether the action sat inside the scope you wrote. Without the inventory there is no answer to give, only an argument to have.
What the logs must hold
For every payment and every external action, log the action taken, the counterparty, the amount where there is one, the timestamp, the task that caused it, and whether a person approved. Record intent and outcome as separate entries so that the two can be compared against each other.
A gap between them is information rather than noise. It is how you notice an endpoint that repriced, an agent making more calls than its task required, or an approval that got recorded without ever being shown to anyone.
Retention should run long enough to reconstruct what your agent did across any window a dispute could reach back into. Pick that span deliberately and write the reasoning next to it, because the reasoning is the part you will be asked about.
Keeping it true
Revise the document when your agent gains a new class of external action, when a price ceiling moves, when a service category is added, or when any action changes between automatic and approval-required.
A scope document describing the agent as it existed six months ago does not establish current authorization. It establishes that someone stopped maintaining the authorization process, which is a weaker position than having written nothing at all.
Sources
https://the-decoder.com/ftc-launches-sweeping-probe-into-openai-anthropic-and-other-ai-labs-over-consumer-protection-concerns/