Visual explainer
The Agent Router Architecture
How placing a fast decision model in front of expensive LLMs saves latency and compute by routing requests deterministically.
When an application defaults to a massive generative model for every request, it incurs heavy latency and high costs—even for trivial queries that could be handled instantly.
The System 1 Intercept
By placing a System 1 router (like Convai's Laya) at the entry point, the system intercepts every request. It reads the state and predicts the category and severity before the LLM is ever invoked.
Routing the Request
The router's typed outputs (Choice, Score) act as traffic controllers. A "Password Reset" goes straight to deterministic code. A "Code Generation" request is routed to an expensive LLM. A "Refund (Angry)" ticket skips the AI entirely and escalates to a human queue.
The Latency Budget Saved
Because the vast majority of incoming requests are simple (e.g., triage, basic support, metadata extraction), they can be resolved or routed correctly in ~35ms. The slow, 3000ms LLM is only invoked when true reasoning or generation is required.
Where It Breaks
If the router is not calibrated (e.g., not trained with RLCD), it will produce overconfident, wrong probabilities. A confident mistake bypasses the human review threshold and incorrectly triggers deterministic code, breaking trust.
The Quick Version
- Monolithic systems are slow and expensive for simple tasks.
- System 1 routers sit at the front to triage requests.
- Typed decisions route traffic to code, humans, or LLMs.
- Latency is minimized by avoiding LLM invocations when unnecessary.
- Calibration is critical to prevent confident misrouting.
What to Read Next
- How A Decision Head WorksVisualizing how a System 1 router replaces autoregressive text generation with a single-pass encoder and typed decision heads.
- Attention MechanismInstead of compressing an entire sequence into one vector, attention lets the decoder dynamically weigh and blend the full input for each output step.