Skip to content
AI360Xpert
Core ML

Visual explainer

The Agent Router Architecture

How placing a fast decision model in front of expensive LLMs saves latency and compute by routing requests deterministically.

A monolithic agent stack forces every request through an expensive, slow LLM.
A monolithic agent stack forces every request through an expensive, slow LLM.

When an application defaults to a massive generative model for every request, it incurs heavy latency and high costs—even for trivial queries that could be handled instantly.

The System 1 Intercept

Placing a fast System 1 router in front intercepts the request in milliseconds.
Placing a fast System 1 router in front intercepts the request in milliseconds.

By placing a System 1 router (like Convai's Laya) at the entry point, the system intercepts every request. It reads the state and predicts the category and severity before the LLM is ever invoked.

Routing the Request

The router directs traffic based on calibrated probabilities to specialized backends.
The router directs traffic based on calibrated probabilities to specialized backends.

The router's typed outputs (Choice, Score) act as traffic controllers. A "Password Reset" goes straight to deterministic code. A "Code Generation" request is routed to an expensive LLM. A "Refund (Angry)" ticket skips the AI entirely and escalates to a human queue.

The Latency Budget Saved

By intercepting simple requests, the router saves massive amounts of latency and compute.
By intercepting simple requests, the router saves massive amounts of latency and compute.

Because the vast majority of incoming requests are simple (e.g., triage, basic support, metadata extraction), they can be resolved or routed correctly in ~35ms. The slow, 3000ms LLM is only invoked when true reasoning or generation is required.

Where It Breaks

Uncalibrated routers will confidently misroute requests if the confidence gate is set improperly.
Uncalibrated routers will confidently misroute requests if the confidence gate is set improperly.

If the router is not calibrated (e.g., not trained with RLCD), it will produce overconfident, wrong probabilities. A confident mistake bypasses the human review threshold and incorrectly triggers deterministic code, breaking trust.

The Quick Version

  • Monolithic systems are slow and expensive for simple tasks.
  • System 1 routers sit at the front to triage requests.
  • Typed decisions route traffic to code, humans, or LLMs.
  • Latency is minimized by avoiding LLM invocations when unnecessary.
  • Calibration is critical to prevent confident misrouting.