Visual explainer
How A Decision Head Works
Visualizing how a System 1 router replaces autoregressive text generation with a single-pass encoder and typed decision heads.
Standard generative models (System 2) are fundamentally probabilistic text generators. When a request needs routing or categorization, parsing generated JSON token by token is a significant bottleneck and a primary source of hallucination.
The Shared Representation
Instead of generating text, a System 1 model (like Laya) uses an encoder backbone to read the entire state (like an email or support ticket) in one pass, building a dense vector representation.
Parallel Decision Heads
Multiple non-autoregressive "heads" sit on top of this shared representation. Each head is trained to answer a specific typed question (a choice, a numeric score, or a boolean) simultaneously in a single forward pass.
Deterministic Typings
Because the model outputs strict types with calibrated probabilities rather than free-form text, there is nothing to parse. The routing happens deterministically in under 50 milliseconds.
Where It Breaks
System 1 models are routers, not reasoners. If the task requires synthesizing new information, generating a response email, or writing code, a decision model is the wrong tool.
The Quick Version
- Autoregressive models generate one text token at a time.
- Decision models use an encoder to read the input state in one pass.
- Typed heads output parallel, non-generative decisions (scores, choices).
- Latency drops dramatically (e.g., 30-50ms) because there's no decoding loop.
- Zero hallucinations occur because no text is ever generated.
What to Read Next
- Attention MechanismInstead of compressing an entire sequence into one vector, attention lets the decoder dynamically weigh and blend the full input for each output step.
- TransformersA visual walkthrough of the Transformer architecture, from self-attention and positional encodings to the full encoder-decoder stack.
- The Agent Router ArchitectureHow placing a fast decision model in front of expensive LLMs saves latency and compute by routing requests deterministically.