Skip to content
AI360Xpert
Beta

Categorical DQN (C51)

Instead of estimating a single expected scalar value for each action, Categorical DQN models the entire probability distribution over future returns across 51 discrete atoms, capturing risk, environmental stochasticity, and multimodality.

The C51 projection operator shifts return atoms by reward and discount before distributing probability mass to adjacent support bins.
The C51 projection operator shifts return atoms by reward and discount before distributing probability mass to adjacent support bins.

Why Does This Exist?

In classical Q-learning, the value function estimates a single scalar expectation:

Q(s,a)=E[∑t=0∞γtRt+1  |  S0=s,A0=a]Q(s, a) = \mathbb{E}\left[ \sum_{t=0}^{\infty} \gamma^t R_{t+1} \;\middle|\; S_0 = s, A_0 = a \right]

Collapsing the future into a single scalar expectation discards critical structural information about the environment:

  1. Risk Blindness: An expected return of +10+10 can represent an invariant deterministic payout of +10+10 with 100%100\% certainty, or a high-stakes coin flip yielding +100+100 (50%50\% probability) versus −80-80 (50%50\% probability). A scalar Q-network treats both actions as strictly identical, making risk-sensitive decision making impossible.
  2. Loss of Multimodality: In stochastic or partially observable domains, transitioning from (s,a)(s, a) frequently branches into distinct probabilistic regimes (e.g., encountering a cooperative opponent versus an aggressive adversary). Averaging these distinct modes into a single number creates an artificial mean that may have zero probability of ever occurring.
  3. Weak Feature Representation Gradients: When deep neural networks minimize scalar Bellman errors (r+γmax⁡Q(s′,a′)−Q(s,a))2(r + \gamma \max Q(s', a') - Q(s, a))^2, gradient signals push the entire feature hierarchy toward a single scalar target. Training on scalar targets frequently causes catastrophic feature drift and policy oscillation.

Introduced by Bellemare, Dabney, and Munos in 2017, Categorical DQN (C51) founded the field of Distributional Reinforcement Learning. Instead of learning the expectation E[Z(s,a)]\mathbb{E}[Z(s, a)], C51 learns the full probability distribution of the return random variable Z(s,a)Z(s, a). By preserving variance, skewness, and multimodality across a 51-atom categorical distribution, C51 provides vastly richer gradient signals that stabilize deep neural network representations.

Think of It Like This

Weather Forecasting: The Entire Rain Distribution vs. The Single Average

Imagine planning an outdoor wedding and consulting two different meteorological forecasting services:

  • The Classical Scalar Forecaster (Standard DQN): The meteorologist announces, "The expected rainfall today is 13mm." This single average tells you almost nothing useful. Is it going to be a gentle, harmless 1mm/hour drizzle all day? Or is there an 80% chance of blue skies and sunshine coupled with a 20% chance of a catastrophic 65mm flash flood? A single expected number completely masks the danger.
  • The Distributional Forecaster (Categorical C51): The meteorologist provides a complete probability histogram over rainfall amounts:
    • 0mm (Sunshine): 70%70\% probability
    • 5mm (Light shower): 10%10\% probability
    • 50mm (Severe storm): 20%20\% probability

With the full probability distribution in hand, you can measure both the expected value (0.7×0+0.1×5+0.2×50=10.5mm0.7 \times 0 + 0.1 \times 5 + 0.2 \times 50 = 10.5\text{mm}) and the tail risk of disaster (20%20\% chance of a severe storm), enabling intelligent, risk-aware decisions.

Where the analogy stops: Weather forecasting is a passive observational task where forecasts do not alter the atmospheric physics. In C51, the agent uses the distribution not only to act (choosing actions by their distribution mean or risk-sensitive metrics), but also iteratively shifts, scales, and projects future distributions through the Bellman operator to bootstrap earlier decisions.

How It Actually Works

The Distributional Bellman Equation and Categorical Projection

Let Z(s,a)Z(s, a) denote the return random variable whose expectation is the action-value function: Q(s,a)=E[Z(s,a)]Q(s, a) = \mathbb{E}[Z(s, a)]. The fundamental insight of Distributional RL is that the Bellman equation applies directly to random variables:

Z(s,a)=DR(s,a)+γZ(S′,A∗)Z(s, a) \stackrel{D}{=} R(s, a) + \gamma Z(S', A^*)

where =D\stackrel{D}{=} denotes equality in distribution, S′∼P(⋅∣s,a)S' \sim \mathcal{P}(\cdot \mid s, a), and A∗=arg⁡max⁡a′∈AE[Z(S′,a′)]A^* = \arg\max_{a' \in \mathcal{A}} \mathbb{E}[Z(S', a')].

+-----------------------------------------------------------------------------------------+|                         C51 DISTRIBUTIONAL REINFORCEMENT LEARNING                       |+-----------------------------------------------------------------------------------------+| Fixed Support Atoms | z_i = V_min + i · Δz       | i ∈ {0, ..., N-1}, N = 51            || Probability Vector  | p(s, a) = softmax(ψ(s, a)) | Σ_i p_i(s, a) = 1.0                  || Bellman Shift       | T_j = clip(R + γ z_j)      | Continuous atoms between support bins|| Projection (Φ)      | m_l += p_j·(u - b_j)       | Distributes mass to floor & ceiling  || Optimization Loss   | L(θ) = - Σ_i m_i log p_i   | Cross-entropy / KL Divergence        |+-----------------------------------------------------------------------------------------+

1. The Categorical Parameterization

C51 approximates the probability distribution of Z(s,a)Z(s, a) using a discrete categorical distribution supported on N=51N = 51 fixed, uniformly spaced atoms:

zi=Vmin⁡+i⋅Δz,i∈{0,1,…,N−1}z_i = V_{\min} + i \cdot \Delta z, \quad i \in \{0, 1, \dots, N - 1\}

where the atom interval Δz\Delta z is defined by:

Δz=Vmax⁡−Vmin⁡N−1\Delta z = \frac{V_{\max} - V_{\min}}{N - 1}

The neural network takes state ss as input and outputs an unnormalized logit vector ψ(s,a)∈RN\boldsymbol{\psi}(s, a) \in \mathbb{R}^N for each action aa. Passing these logits through a softmax layer yields the probability distribution p(s,a)\mathbf{p}(s, a):

pi(s,a)=exp⁡(ψi(s,a))∑k=0N−1exp⁡(ψk(s,a))p_i(s, a) = \frac{\exp(\psi_i(s, a))}{\sum_{k=0}^{N-1} \exp(\psi_k(s, a))}

The expected Q-value is computed as the dot product between the fixed atoms and their assigned probabilities:

Q(s,a)=∑i=0N−1zi pi(s,a)Q(s, a) = \sum_{i=0}^{N-1} z_i \, p_i(s, a)

The greedy action selected by the agent is a∗=arg⁡max⁡aQ(s,a)a^* = \arg\max_{a} Q(s, a).

2. The Distributional Bellman Projection Operator Φ\Phi

When evaluating a transition (s,a,r,s′)(s, a, r, s'), applying the Bellman operator to atom zjz_j shifts and shrinks its coordinate:

T^zj=r+γzj\hat{\mathcal{T}} z_j = r + \gamma z_j

Because rr and γ\gamma can take arbitrary real values, the shifted coordinate T^zj\hat{\mathcal{T}} z_j will generally not align with any of the predefined discrete atoms {z0,…,zN−1}\{z_0, \dots, z_{N-1}\}, and may fall outside [Vmin⁡,Vmax⁡][V_{\min}, V_{\max}].

To map the shifted distribution back onto the original support, C51 introduces the projection operator Φ\Phi:

  1. Clamping to Support Bounds: T^j=clip(r+γzj, Vmin⁡, Vmax⁡)\hat{T}_j = \text{clip}\left( r + \gamma z_j, \, V_{\min}, \, V_{\max} \right)
  2. Continuous Atom Coordinate: bj=T^j−Vmin⁡Δz∈[0,N−1]b_j = \frac{\hat{T}_j - V_{\min}}{\Delta z} \in [0, N - 1]
  3. Neighboring Integer Atoms: l=⌊bj⌋,u=⌈bj⌉l = \lfloor b_j \rfloor, \quad u = \lceil b_j \rceil
  4. Linear Mass Interpolation: The probability mass pj(s′,a∗)p_j(s', a^*) from atom jj of the target distribution is apportioned to neighboring integer bins ll and uu inversely proportional to distance: ml←ml+pj(s′,a∗)⋅(u−bj)m_l \leftarrow m_l + p_j(s', a^*) \cdot (u - b_j) mu←mu+pj(s′,a∗)⋅(bj−l)m_u \leftarrow m_u + p_j(s', a^*) \cdot (b_j - l) If bjb_j lands exactly on an integer atom (l=ul = u), the entire probability mass pj(s′,a∗)p_j(s', a^*) is assigned directly to mlm_l.

3. Cross-Entropy Loss

Because both the projected target m\mathbf{m} and the current network prediction p(s,a;θ)\mathbf{p}(s, a; \boldsymbol{\theta}) share the identical support {z0,…,zN−1}\{z_0, \dots, z_{N-1}\}, the parameter update minimizes the Kullback-Leibler (KL) divergence, equivalent to the cross-entropy loss:

L(θ)=DKL(m∥p(s,a;θ))=−∑i=0N−1milog⁡pi(s,a;θ)\mathcal{L}(\boldsymbol{\theta}) = D_{\text{KL}}(\mathbf{m} \parallel \mathbf{p}(s, a; \boldsymbol{\theta})) = - \sum_{i=0}^{N-1} m_i \log p_i(s, a; \boldsymbol{\theta})

Worked numerical example

Let us trace the complete C51 projection operator by hand on a manageable toy support with N=5N = 5 atoms.

1. Setup Parameters

  • Boundaries: Vmin⁡=−2.0,Vmax⁡=2.0,N=5V_{\min} = -2.0, \quad V_{\max} = 2.0, \quad N = 5.
  • Step size: Δz=2.0−(−2.0)5−1=4.04=1.0\Delta z = \frac{2.0 - (-2.0)}{5 - 1} = \frac{4.0}{4} = 1.0
  • Support atoms: z=[z0,z1,z2,z3,z4]=[−2.0,−1.0,0.0,1.0,2.0]\mathbf{z} = [z_0, z_1, z_2, z_3, z_4] = [-2.0, -1.0, 0.0, 1.0, 2.0]
  • Next-state target distribution at greedy action a∗a^*: p(s′,a∗)=[0.10,0.20,0.40,0.20,0.10]\mathbf{p}(s', a^*) = [0.10, 0.20, 0.40, 0.20, 0.10] Notice that E[Z(s′,a∗)]=(−2)(0.1)+(−1)(0.2)+(0)(0.4)+(1)(0.2)+(2)(0.1)=0.00\mathbb{E}[Z(s', a^*)] = (-2)(0.1) + (-1)(0.2) + (0)(0.4) + (1)(0.2) + (2)(0.1) = 0.00.
  • Observed transition: Reward R=0.5R = 0.5, discount factor γ=0.9\gamma = 0.9.

2. Shifting and Projecting Each Atom

Initialize the projected probability mass vector: m=[0.0,0.0,0.0,0.0,0.0]⊤\mathbf{m} = [0.0, 0.0, 0.0, 0.0, 0.0]^\top.

  • Atom j=0j = 0 (z0=−2.0z_0 = -2.0, mass p0=0.10p_0 = 0.10):

    • Shifted location: T^z0=0.5+0.9(−2.0)=0.5−1.8=−1.3\hat{\mathcal{T}} z_0 = 0.5 + 0.9(-2.0) = 0.5 - 1.8 = -1.3.
    • Continuous coordinate: b0=−1.3−(−2.0)1.0=0.70b_0 = \frac{-1.3 - (-2.0)}{1.0} = 0.70.
    • Neighbors: l=0l = 0, u=1u = 1.
    • Allocation weights: u−b0=1−0.70=0.30u - b_0 = 1 - 0.70 = 0.30; b0−l=0.70−0=0.70b_0 - l = 0.70 - 0 = 0.70.
    • Mass contributions:
      • m0+=0.10×0.30=0.03m_0 \mathrel{+}= 0.10 \times 0.30 = 0.03
      • m1+=0.10×0.70=0.07m_1 \mathrel{+}= 0.10 \times 0.70 = 0.07
  • Atom j=1j = 1 (z1=−1.0z_1 = -1.0, mass p1=0.20p_1 = 0.20):

    • Shifted location: T^z1=0.5+0.9(−1.0)=0.5−0.9=−0.4\hat{\mathcal{T}} z_1 = 0.5 + 0.9(-1.0) = 0.5 - 0.9 = -0.4.
    • Continuous coordinate: b1=−0.4−(−2.0)1.0=1.60b_1 = \frac{-0.4 - (-2.0)}{1.0} = 1.60.
    • Neighbors: l=1l = 1, u=2u = 2.
    • Allocation weights: u−b1=2−1.60=0.40u - b_1 = 2 - 1.60 = 0.40; b1−l=1.60−1=0.60b_1 - l = 1.60 - 1 = 0.60.
    • Mass contributions:
      • m1+=0.20×0.40=0.08m_1 \mathrel{+}= 0.20 \times 0.40 = 0.08
      • m2+=0.20×0.60=0.12m_2 \mathrel{+}= 0.20 \times 0.60 = 0.12
  • Atom j=2j = 2 (z2=0.0z_2 = 0.0, mass p2=0.40p_2 = 0.40):

    • Shifted location: T^z2=0.5+0.9(0.0)=0.5\hat{\mathcal{T}} z_2 = 0.5 + 0.9(0.0) = 0.5.
    • Continuous coordinate: b2=0.5−(−2.0)1.0=2.50b_2 = \frac{0.5 - (-2.0)}{1.0} = 2.50.
    • Neighbors: l=2l = 2, u=3u = 3.
    • Allocation weights: u−b2=3−2.50=0.50u - b_2 = 3 - 2.50 = 0.50; b2−l=2.50−2=0.50b_2 - l = 2.50 - 2 = 0.50.
    • Mass contributions:
      • m2+=0.40×0.50=0.20m_2 \mathrel{+}= 0.40 \times 0.50 = 0.20
      • m3+=0.40×0.50=0.20m_3 \mathrel{+}= 0.40 \times 0.50 = 0.20
  • Atom j=3j = 3 (z3=1.0z_3 = 1.0, mass p3=0.20p_3 = 0.20):

    • Shifted location: T^z3=0.5+0.9(1.0)=1.4\hat{\mathcal{T}} z_3 = 0.5 + 0.9(1.0) = 1.4.
    • Continuous coordinate: b3=1.4−(−2.0)1.0=3.40b_3 = \frac{1.4 - (-2.0)}{1.0} = 3.40.
    • Neighbors: l=3l = 3, u=4u = 4.
    • Allocation weights: u−b3=4−3.40=0.60u - b_3 = 4 - 3.40 = 0.60; b3−l=3.40−3=0.40b_3 - l = 3.40 - 3 = 0.40.
    • Mass contributions:
      • m3+=0.20×0.60=0.12m_3 \mathrel{+}= 0.20 \times 0.60 = 0.12
      • m4+=0.20×0.40=0.08m_4 \mathrel{+}= 0.20 \times 0.40 = 0.08
  • Atom j=4j = 4 (z4=2.0z_4 = 2.0, mass p4=0.10p_4 = 0.10):

    • Shifted location: T^z4=0.5+0.9(2.0)=2.3\hat{\mathcal{T}} z_4 = 0.5 + 0.9(2.0) = 2.3.
    • Clamping: Because 2.3>Vmax⁡2.3 > V_{\max}, it is clamped to T^4=2.0\hat{T}_4 = 2.0.
    • Continuous coordinate: b4=2.0−(−2.0)1.0=4.00b_4 = \frac{2.0 - (-2.0)}{1.0} = 4.00.
    • Neighbors: l=4l = 4, u=4u = 4 (l=ul = u).
    • Allocation weights: Entire mass allocated to m4m_4.
    • Mass contribution:
      • m4+=0.10×1.00=0.10m_4 \mathrel{+}= 0.10 \times 1.00 = 0.10

3. Summing Total Projected Distribution

Aggregating contributions across all atoms:

m0=0.03m_0 = 0.03 m1=0.07+0.08=0.15m_1 = 0.07 + 0.08 = 0.15 m2=0.12+0.20=0.32m_2 = 0.12 + 0.20 = 0.32 m3=0.20+0.12=0.32m_3 = 0.20 + 0.12 = 0.32 m4=0.08+0.10=0.18m_4 = 0.08 + 0.10 = 0.18

Check normalization:

∑i=04mi=0.03+0.15+0.32+0.32+0.18=1.00\sum_{i=0}^4 m_i = 0.03 + 0.15 + 0.32 + 0.32 + 0.18 = 1.00

The expected value of the projected distribution is:

E[m]=(−2)(0.03)+(−1)(0.15)+(0)(0.32)+(1)(0.32)+(2)(0.18)=−0.06−0.15+0.00+0.32+0.36=0.47\mathbb{E}[\mathbf{m}] = (-2)(0.03) + (-1)(0.15) + (0)(0.32) + (1)(0.32) + (2)(0.18) = -0.06 - 0.15 + 0.00 + 0.32 + 0.36 = 0.47

(Note: The unclamped mathematical expectation was R+γE[Z]=0.5+0.9(0)=0.50R + \gamma \mathbb{E}[Z] = 0.5 + 0.9(0) = 0.50. The 0.030.03 difference arises precisely because the upper tail atom z4=2.0z_4 = 2.0 shifted beyond Vmax⁡V_{\max} and was clamped, illustrating the boundary clipping effect).

Code

The following self-contained Python script implements the full C51 projection operator Φ\Phi and cross-entropy loss function, verifying both the 5-atom worked numerical example and a full 51-atom batch scenario with automated assertions.

from typing import Tupleimport numpy as np

def project_distribution(    next_dist: np.ndarray,    rewards: np.ndarray,    dones: np.ndarray,    gamma: float = 0.99,    v_min: float = -10.0,    v_max: float = 10.0,    n_atoms: int = 51,) -> np.ndarray:    """Project next-state return distributions onto fixed support atoms via C51 operator Phi.
    Args:        next_dist: Probabilities for greedy next action, shape (batch_size, n_atoms).        rewards: Transition rewards, shape (batch_size,).        dones: Terminal flags, shape (batch_size,).        gamma: Discount factor.        v_min: Lower bound of the support.        v_max: Upper bound of the support.        n_atoms: Number of categorical atoms.
    Returns:        Projected probability distribution m of shape (batch_size, n_atoms).    """    batch_size = rewards.shape[0]    delta_z = (v_max - v_min) / (n_atoms - 1)    atoms = np.linspace(v_min, v_max, n_atoms)
    # Shifted atoms: T_j = r + gamma * z_j (or r if terminal)    tz = np.where(        dones[:, None],        rewards[:, None],        rewards[:, None] + gamma * atoms[None, :],    )    # Clamp to [v_min, v_max]    tz = np.clip(tz, v_min, v_max)
    # Continuous coordinate on the support grid: bj in [0, n_atoms - 1]    b = (tz - v_min) / delta_z    l = np.floor(b).astype(np.int64)    u = np.ceil(b).astype(np.int64)
    m = np.zeros((batch_size, n_atoms), dtype=np.float64)
    # Distribute probability mass to adjacent integer bins    for i in range(batch_size):        for j in range(n_atoms):            p_j = next_dist[i, j]            lj, uj, bj = l[i, j], u[i, j], b[i, j]            if lj == uj:                m[i, lj] += p_j            else:                m[i, lj] += p_j * (uj - bj)                m[i, uj] += p_j * (bj - lj)
    return m

def categorical_cross_entropy(    predicted_dist: np.ndarray, target_dist: np.ndarray, eps: float = 1e-8) -> float:    """Compute categorical cross-entropy loss: L = - sum(m * log(p))."""    p_safe = np.clip(predicted_dist, eps, 1.0)    ce = -np.sum(target_dist * np.log(p_safe), axis=-1)    return float(np.mean(ce))

if __name__ == "__main__":    # 1. Verify 5-Atom Worked Numerical Example    p_next_toy = np.array([[0.10, 0.20, 0.40, 0.20, 0.10]])    r_toy = np.array([0.5])    d_toy = np.array([False])
    m_toy = project_distribution(        p_next_toy, r_toy, d_toy, gamma=0.90, v_min=-2.0, v_max=2.0, n_atoms=5    )
    print("=== C51 5-Atom Worked Example ===")    print("Projected m:", np.round(m_toy[0], 4))    toy_atoms = np.linspace(-2.0, 2.0, 5)    toy_expected = float(np.sum(m_toy[0] * toy_atoms))    print(f"Projected Mean E[m]: {toy_expected:.4f}")
    # Automated assertions matching worked example    assert np.allclose(m_toy[0], [0.03, 0.15, 0.32, 0.32, 0.18], atol=1e-4)    assert np.isclose(toy_expected, 0.47, atol=1e-4)    assert np.isclose(np.sum(m_toy[0]), 1.0, atol=1e-6)
    # 2. Verify Full 51-Atom Batch Simulation    rng = np.random.default_rng(42)    B, N = 4, 51
    logits_pred = rng.standard_normal((B, N))    pred_dist = np.exp(logits_pred) / np.sum(        np.exp(logits_pred), axis=-1, keepdims=True    )
    logits_next = rng.standard_normal((B, N))    next_dist = np.exp(logits_next) / np.sum(        np.exp(logits_next), axis=-1, keepdims=True    )
    batch_rewards = np.array([1.0, -2.5, 0.0, 5.0])    batch_dones = np.array([False, False, True, False])
    m_batch = project_distribution(        next_dist,        batch_rewards,        batch_dones,        gamma=0.99,        v_min=-10.0,        v_max=10.0,        n_atoms=N,    )
    loss = categorical_cross_entropy(pred_dist, m_batch)
    print("\n=== C51 51-Atom Batch Results ===")    print(f"Batch size: {B}, Support: [-10.0, 10.0], Atoms: {N}")    for b in range(B):        sum_m = np.sum(m_batch[b])        atoms_51 = np.linspace(-10.0, 10.0, N)        mean_val = np.sum(m_batch[b] * atoms_51)        print(            f"Sample {b}: sum(m) = {sum_m:.6f}, Mean Return = {mean_val:.4f}"        )        assert np.isclose(sum_m, 1.0, atol=1e-6)
    print(f"\nCategorical Cross-Entropy Loss: {loss:.6f}")    print("All assertions passed successfully!")
# Expected Output:# === C51 5-Atom Worked Example ===# Projected m: [0.03 0.15 0.32 0.32 0.18]# Projected Mean E[m]: 0.4700## === C51 51-Atom Batch Results ===# Batch size: 4, Support: [-10.0, 10.0], Atoms: 51# Sample 0: sum(m) = 1.000000, Mean Return = 3.0438# Sample 1: sum(m) = 1.000000, Mean Return = -2.4684# Sample 2: sum(m) = 1.000000, Mean Return = 0.0000# Sample 3: sum(m) = 1.000000, Mean Return = 3.8061## Categorical Cross-Entropy Loss: 4.488480# All assertions passed successfully!

Watch Out For

Support Boundary Clamping and Resolution Trade-Off

C51's greatest architectural vulnerability is its reliance on a fixed, static support [Vmin⁡,Vmax⁡][V_{\min}, V_{\max}].

The Failure Mode:

  • Boundary Truncation (Under-bounding): If you choose Vmax⁡=10V_{\max} = 10 in an environment where true discounted returns reach +100+100, any return higher than +10+10 is forcibly clamped into the 51st atom z50z_{50}. The entire right tail of the return distribution is compressed into a single spike at 1010, severely underestimating high-reward trajectories and biasing policy choices.
  • Resolution Dilution (Over-bounding): Conversely, setting [Vmin⁡,Vmax⁡][V_{\min}, V_{\max}] unnecessarily broad (e.g., [−1000,1000][-1000, 1000]) means the fixed 51 atoms are spaced far apart (Δz≈40\Delta z \approx 40). Small differences between +1+1 and +5+5 are lost within a single bin, wiping out the fine-grained value comparisons needed to select optimal actions.

The Fix:

  1. Carefully calibrate [Vmin⁡,Vmax⁡][V_{\min}, V_{\max}] using prior domain knowledge or normalize episodic rewards.
  2. Upgrade to Quantile Regression DQN (QR-DQN): Rather than fixing atom locations and learning probabilities, QR-DQN fixes probabilities (pi=1/Np_i = 1/N) and allows the network to predict dynamic, adaptive return quantile locations, completely eliminating the [Vmin⁡,Vmax⁡][V_{\min}, V_{\max}] bounding box restriction.

The Quick Version

  • Return Distribution: Instead of predicting a scalar average E[Gt]\mathbb{E}[G_t], C51 models the full probability distribution of future returns Z(s,a)Z(s, a), preserving variance, multimodality, and tail risk.
  • Categorical Parameterization: The return distribution is represented by N=51N=51 discrete, equidistant atoms {z0,…,z50}\{z_0, \dots, z_{50}\} bounded between Vmin⁡V_{\min} and Vmax⁡V_{\max}, with probabilities generated via a neural softmax.
  • Projection Operator Φ\Phi: Applying the Bellman shift (r+γzjr + \gamma z_j) produces non-integer coordinates that fall between fixed support atoms; the C51 projection operator linearly distributes probability mass to the adjacent floor and ceiling bins.
  • Optimization Objective: The online network is trained to match the projected Bellman target distribution by minimizing the categorical cross-entropy loss (Kullback-Leibler divergence).