How Your Agent Actually Spends Money: Mechanics, Records, and Controls
When your agent makes a call to a paid endpoint, the payment happens in the same HTTP exchange as the request. You are not approving individual transactions. Understanding the mechanics: the wallet, the trail, and the counterparty, is the prerequisite to building controls that work.
· 1083 words
When your agent makes a call to a paid endpoint, the payment happens in the same HTTP exchange as the request. You are not approving individual transactions. Your agent is spending from a wallet you funded, to payment addresses associated with services you may never have reviewed in detail. Understanding what actually happens in that exchange is the prerequisite to building controls that hold.
The payment model determines when you are in the loop
The payment model your agent uses depends on what it is buying.
For per-call microservices, the agent holds a pre-funded wallet and pays per request. No confirmation. No human prompt. The call fires, the payment settles, the response comes back. The protocol that wires this into HTTP is x402: a 402 response from the endpoint triggers a payment challenge, the agent pays from its wallet, and the request retries with proof of payment. The median paid agent endpoint in the current market charges $0.01 per call. At that price, the individual transaction is not the risk. The risks are volume and drift: an agent making a call in a loop, an endpoint that reprices without notice, or an agent that funds a service heavily without your awareness.
For higher-value purchases, the model shifts to human-authorized checkout. The agent handles the mechanics of the transaction, assembling the order and verifying the details, but a human confirms before the final step fires. Shopify's WebMCP checkout tools work this way: the agent can read and modify the checkout, but the submit step requires explicit buyer authorization. This model is explicit about when the human is in the loop. The per-call model is not.
Knowing which model applies to a given purchase is the first decision. A per-call payment at $0.01 requires no approval workflow. A checkout that ships a physical product to a physical address requires one. Conflating them is where both failure modes live: unnecessary friction on routine calls, or charges that happen without adequate notice.
The trail: what records actually exist
Two records should exist for any payment your agent makes. The first is the agent's own log of what it intended to pay. The second is the on-chain settlement record, which is what actually happened.
Of the services indexed in the paid endpoint market, 2,110 hosts have on-chain payment records that can be verified against a ledger. For a payment to a host in this group, you can confirm independently that the payment occurred, to which address, and at what amount. For a payment to a host outside this group, your verification options are different: you are relying on the agent's self-report and on whatever logging the endpoint itself provides.
This is not a statement about trust. It is a statement about audit capability. An on-chain record is a check you can run yourself, independently of anything the agent or the endpoint tells you. Plan your verification approach based on whether that check is available.
Who you are actually paying
A payment address is not the same as a vetted counterparty. Registry listings that include a payment address can be stale, can be placeholder entries set up but never activated, or can belong to operators who changed their service without updating their listing.
The clearest proxy for counterparty substance is corroboration: whether the same host appears in more than one independent registry. Of the hosts currently indexed, 1,304 appear in two or more independent registries. That is evidence of existence, not evidence of quality: two independent registry operators assessed the host as worth listing. For an endpoint you are about to authorize for wallet access, that distinction matters.
For hosts that appear in only one registry: the listing is your only evidence. Supplement it with a liveness check (does the endpoint respond to a handshake?), a scope check (does it declare what it charges for what operations?), and a price estimate before committing to a task. None of these are substitutes for corroboration, but they are checks you can run before the first payment.
Three controls to set before your agent spends
Wallet isolation. Your agent's payment wallet should hold only the funds you have authorized for agent spending. Not your main treasury, not a shared pool. Fund the wallet in increments. This limits the exposure if the agent makes an unexpected payment or if an endpoint charges more than you anticipated. The question is not whether your agent will behave correctly. It is what your downside looks like when something goes wrong.
Budget enforcement at the infrastructure level. A budget cap in your agent's code is a reasonable control. It is not sufficient on its own, because your agent's code can be wrong. The more durable control is a ceiling enforced at the wallet level: the wallet will not authorize transactions beyond a defined daily or weekly cap, regardless of what the agent requests. This makes the cap a property of the payment infrastructure, not of the agent logic. Agent logic can be overridden. Wallet-level caps cannot be overridden by the agent.
Address allowlisting. If your agent is authorized to pay a specific set of services, the wallet should reflect that. A wallet that pays any address the agent requests is a wallet that pays any address anyone can convince your agent to call. Address allowlisting converts the decision from the agent decides where to send money to the agent chooses from a set you approved. This control closes the direct path from prompt injection or memory tampering to unauthorized payments. It is the single highest-value control available at the wallet layer.
What the logs should capture
For each payment, your logs should record: the endpoint being called, the payment address the agent sent to, the amount, the timestamp, and the task that triggered the call.
Log the intent before the call. Log the settlement after it. Compare them.
A mismatch between intended and settled payment is not always an error to suppress. It can be the signal that an endpoint repriced, that the agent made more calls than the task required, or that something in the payment path did not work as intended. The agents whose operators treat payment log mismatches as information rather than noise are the ones who discover endpoint repricing before it becomes a budget problem, and who catch unexpected call patterns before they become a security problem.
The mechanics are not complicated. The defaults are not safe. The gap between those two facts is where the work is.