Skip to content
AI360Xpert
Core ML

Visual explainer

How A Decision Head Works

Visualizing how a System 1 router replaces autoregressive text generation with a single-pass encoder and typed decision heads.

Generative LLMs must produce outputs token by token, making structured JSON parsing slow and error-prone.
Generative LLMs must produce outputs token by token, making structured JSON parsing slow and error-prone.

Standard generative models (System 2) are fundamentally probabilistic text generators. When a request needs routing or categorization, parsing generated JSON token by token is a significant bottleneck and a primary source of hallucination.

The Shared Representation

A System 1 model uses an encoder to read the state once and build a shared, dense representation.
A System 1 model uses an encoder to read the state once and build a shared, dense representation.

Instead of generating text, a System 1 model (like Laya) uses an encoder backbone to read the entire state (like an email or support ticket) in one pass, building a dense vector representation.

Parallel Decision Heads

Multiple typed decision heads run in parallel over that shared representation to produce distinct answers.
Multiple typed decision heads run in parallel over that shared representation to produce distinct answers.

Multiple non-autoregressive "heads" sit on top of this shared representation. Each head is trained to answer a specific typed question (a choice, a numeric score, or a boolean) simultaneously in a single forward pass.

Deterministic Typings

The result is instantaneous, deterministic answers without the need to parse generated JSON.
The result is instantaneous, deterministic answers without the need to parse generated JSON.

Because the model outputs strict types with calibrated probabilities rather than free-form text, there is nothing to parse. The routing happens deterministically in under 50 milliseconds.

Where It Breaks

A decision model cannot generate text. If you need a summary or translation, you still need an LLM.
A decision model cannot generate text. If you need a summary or translation, you still need an LLM.

System 1 models are routers, not reasoners. If the task requires synthesizing new information, generating a response email, or writing code, a decision model is the wrong tool.

The Quick Version

  • Autoregressive models generate one text token at a time.
  • Decision models use an encoder to read the input state in one pass.
  • Typed heads output parallel, non-generative decisions (scores, choices).
  • Latency drops dramatically (e.g., 30-50ms) because there's no decoding loop.
  • Zero hallucinations occur because no text is ever generated.