AgentIndex · traderszone

AgentIndex · Guides

Routing as a Service: Auditing Your Orchestrator's Dispatch Step

An orchestrator's routing step runs on every task, which makes it one of the highest-frequency operations in a multi-agent system. The per-call microservice infrastructure already exists for it. Here is how to decide whether your routing step belongs there.

· 781 words

An orchestrator's routing step runs on every task. Every task routes. That makes routing one of the highest-frequency operations in a multi-agent system and a step worth examining when the cost of per-call services drops.

The per-call microservice market has been forming. The paid endpoint index records a $0.01 median charge per call across active services. Of listed paid endpoints, 15.8% have more than one distinct paying buyer, which is the measure of real commercial activity versus services with a single operator calling themselves. These two facts describe an infrastructure layer that exists and is being bought.

The Decisions API, which OpenAI launched at DevDay on Tuesday, enters that infrastructure layer as a specialized product: a per-call endpoint that returns an answer from a predefined set. Its existence reflects a design premise worth testing against your own orchestrator: the routing step may be the part of your multi-agent workflow best served by a purpose-built per-call service.

The routing step's cost profile

Routing sits on the critical path of every task. A routing call on every task means routing is a per-task cost, not an occasional one. At the per-call pricing that the current microservice market uses, routing is predictable and cheap. At the price of a full general inference model, the same call volume carries the cost of inference you are not using.

Routing is the one step in a multi-agent workflow that qualifies for a constrained model, because routing has a defined answer space: send this task to the code agent, the search agent, or the retrieval agent. A fixed answer space is exactly what a decision endpoint is designed for.

What to check before committing to a decision endpoint

The endpoint must offer three things that general inference does not guarantee by default.

Defined answer validation. The endpoint accepts your answer set at registration and returns only answers from that set. An endpoint that can return freeform text is not a decision endpoint. The distinction matters because downstream handlers in a multi-agent workflow are written to handle specific routing outcomes. An unexpected output breaks the handler.

Latency transparency. Routing sits on the critical path of every task. A routing call that adds latency adds it on every task your orchestrator runs. If the decision endpoint does not publish a latency target, test median and tail latency under realistic load before building your routing logic on it.

Per-call pricing. Decision endpoints in the current market price per call. Before committing, verify the pricing is per call and not per token. A token-based pricing model applies to both the context you send and the answer you receive. For routing calls with large context payloads, per-token pricing can cost more than a general inference call.

Auditing your orchestrator for outsource candidates

Not every step qualifies. The constraints are real: fixed answer space, high frequency, latency sensitivity, and context compact enough for the pricing to work. Run through your orchestrator's dispatch logic and mark every step that:

  • Returns one of a predefined set of outputs
  • Runs on every task or on a large proportion of tasks
  • Does not require reasoning about situations not covered by the predefined options

Steps that meet all three criteria are candidates. Steps that require open-ended reasoning, or that run infrequently, are not worth the integration cost.

The routing step is the canonical example. Intent classification, content routing, binary policy enforcement, and priority queuing are others. OpenAI's stated use cases for the Decisions API match this list, which is the signal that the product reflects observed demand rather than speculation about what developers might want.

What to do when the answer space is not yet fixed

Some orchestrators have not defined their routing answer space. They route dynamically, letting the model decide which downstream agent to use based on open-ended reasoning. That is a valid design for exploratory workflows, but it is not a design you can move to a decision endpoint without first pinning the answer space.

If your routing logic is dynamic, the decision-endpoint question is premature. The productive prior step is to instrument which downstream agents actually get called for which task types, look for the cases where the routing decision is predictable, and extract those into a fixed routing policy. When that fixed policy covers a large enough fraction of your task volume, it becomes worth moving to a decision endpoint.

The per-call infrastructure layer already exists and has proven demand. A purpose-built service for one of its core use cases entered the market this week with a competing product already available. The question is not whether this is viable. It is whether your specific routing step qualifies.

This came from the index.

AgentIndex probes agentic endpoints rather than repeating their listings. Browse what we measured, or point your agent at it.