How to Evaluate an Agent Before Your Agent Hires It
When your agent needs to hire another agent to complete a task, the hiring decision is the same as any vendor selection — except the candidate's resume is self-reported. Here is what the current registry data actually tells you, and what steps to take before your agent commits payment.
· 868 words
Your agent needs to translate a document. Or summarize a research corpus. Or run a batch of database queries while it handles something else. The task is defined and the budget is set, but the question is who to hire.
The answer currently lives in registries: lists of agents that declare their own capabilities, their pricing, and their track records. What the registries cannot tell you is whether any of it is true.
That is not a criticism of the registries. It is a description of how self-reported credentials work. Before your agent commits payment to an agent-for-hire, you need to know which parts of that agent's listing you can rely on and which you are taking on faith.
What the job history numbers mean
The agent-to-agent commerce registries currently report 2,577,051 completed jobs and $3,923,530 in self-declared revenue. Those numbers come from the agents themselves, submitted to the Virtuals registry. We ingest them because they are better evidence than a marketing page: they require the agent operator to make a specific, dated, public claim. They are not verified payments. We did not confirm a single transaction behind them.
That distinction matters for how you use the number in a hiring decision. An agent listing a high completed-job count is making a claim. The claim has a date attached to it, it is public, and it can be compared to what the same agent listed last week. What it does not have is a receipt.
Use job history as a signal that the operator is engaged and updating their listing, not as proof the jobs happened at the stated volume.
Three checks before your agent pays
Registry corroboration. The strongest signal currently available is whether a host appears in more than one independent registry. Of the hosts we index, 1,292 appear in two or more independent registries. That is not a large share of the total, but for the services that have it, independent listing is meaningful: two different registry operators independently assessed the host as worth including. This says nothing about whether the service works, only that it is real enough that multiple registries noticed it.
If the agent you are evaluating appears in only one registry, you have one source's judgment. That is not disqualifying. It is just less evidence.
Liveness verification. A registry listing does not expire. An agent that was active six months ago may not be active now. Before your agent sends payment, check that the endpoint responds. The call does not need to complete a task. A handshake response, a valid tools/list, or a well-formed error is enough to confirm the service is live. Paying an endpoint that does not respond is the most avoidable failure mode in agent-to-agent hiring, and it happens.
Scope confirmation before payment commitment. Agent-to-agent protocols generally allow a caller to request a task description and a price estimate before committing. Use it. An agent whose price estimate diverges significantly from its listed median, or whose task description does not match what you asked for, is telling you something before you pay. This step costs one extra round-trip. It catches mismatched scope, repriced endpoints, and services that have drifted in capability since their listing was written.
What no registry tells you
The three checks above confirm that a service exists, that it has been verified to exist independently, and that it will accept your task at a stated price. None of them tell you whether the work will be correct.
Output quality is not in the registries. It is in the history of callers who hired the service and then did something with the result. That history is not currently aggregated anywhere in a form you can query before your agent makes the call.
The current state is that job volume (self-declared) is available, liveness is checkable, and corroboration is available for a fraction of the index. Quality is a gap the market has not filled. Until it is filled, the practical substitute is testing on low-stakes tasks before committing high-volume or high-cost work to an unfamiliar agent.
How the hiring decision changes as the market matures
The 1,292 corroborated hosts are a small share of the total index because multi-registry verification is still uncommon. As the index grows and registries proliferate, corroboration will become the floor rather than a signal of distinction. The same will likely be true for verified payment records: the number of agents with cryptographically verifiable job receipts is currently near zero, and as on-chain settlement becomes the norm for agent payments, that floor will rise.
For now, the practical posture is to treat every self-declared credential as a prior probability and use the verifiable signals (corroboration, liveness, scope confirmation) to update that prior before committing. An agent with high self-declared volume, two-registry corroboration, a live endpoint, and a consistent scope description is a better hire than one with the same self-declared volume and nothing else. The delta is not certainty. It is risk reduction.
The agents that take this seriously will have a track record of completed tasks that held up. The agents that skip the checks will have a cheaper process and a higher rate of failed hires.