Skip to content
AI360Xpert
Beta

Actor-Critic Architecture

Actor-critic splits reinforcement learning into two cooperating networks: an Actor that selects actions and a Critic that grades those choices step-by-step using temporal difference errors.

Actor-Critic dual network architecture where a shared representation branches into a policy head (Actor) and a value baseline head (Critic), coordinated via TD advantage.
Actor-Critic dual network architecture where a shared representation branches into a policy head (Actor) and a value baseline head (Critic), coordinated via TD advantage.

Why Does This Exist?

Pure policy gradient methods like REINFORCE suffer from excessive variance because they evaluate actions using complete trajectory returns Gt=∑γkRt+k+1G_t = \sum \gamma^k R_{t+k+1}. In long-horizon tasks, small variations in environment randomness thousands of steps downstream corrupt the gradient of an action taken at step zero. Conversely, pure value-based methods like Q-learning bootstrap efficiently on single-step transitions but cannot natively parameterize continuous action spaces or learn stochastic policies.

The Actor-Critic architecture unites both families into a single hybrid system. An Actor parameterizes the policy πθ(a∣s)\pi_\theta(a \mid s), while a Critic parameterizes a value function Vϕ(s)V_\phi(s).

Instead of waiting for an entire episode to finish, the Actor updates after every transition (or over small nn-step rollout windows) using temporal difference (TD) errors provided by the Critic. By replacing the noisy Monte Carlo return GtG_t with a 1-step bootstrapped estimate Rt+1+γVϕ(St+1)R_{t+1} + \gamma V_\phi(S_{t+1}), Actor-Critic methods drastically slash variance while retaining the ability to optimize arbitrary continuous or discrete policies.

Think of It Like This

A stage actor rehearsing with a director giving instant notes

Imagine a theater performer (the Actor) rehearsing lines for an upcoming play.

Under a pure Monte Carlo approach (REINFORCE), the performer would deliver all five acts of the play from start to finish, wait until opening night to see the final newspaper review, and only then try to recall which specific facial expressions in Act 1 contributed to the critic's overall verdict. Learning this way is excruciatingly slow and full of second-guessing.

Under an Actor-Critic setup, a professional director (the Critic) sits in the front row during rehearsals. The performer delivers a single line (takes action AtA_t). Immediately, the director interrupts: "That inflection was sharper than usual, keep that" (positive TD error), or "That tone fell flat compared to what this scene requires" (negative TD error).

The performer immediately adjusts their line delivery (updates policy parameters θ\theta) based on the director's localized feedback. Meanwhile, the director continually refines their own mental standard of how much applause each scene is worth (updates value parameters ϕ\phi).

How It Actually Works

Advantage Estimation and Dual Objective Updates

In modern deep Actor-Critic architectures, the Actor and Critic typically share a common convolutional or transformer representation backbone ht=fbase(St;θbase)h_t = f_{\text{base}}(S_t; \theta_{\text{base}}), branching into two output heads:

  1. Actor Head: Outputs action distribution parameters πθ(a∣St)\pi_\theta(a \mid S_t) (e.g., categorical logits or Gaussian mean and standard deviation).
  2. Critic Head: Outputs a scalar state-value baseline Vϕ(St)∈RV_\phi(S_t) \in \mathbb{R}.

When the agent executes action At∼πθ(⋅∣St)A_t \sim \pi_\theta(\cdot \mid S_t), receives reward Rt+1R_{t+1}, and transitions to St+1S_{t+1}, the Critic computes the 1-step Temporal Difference Advantage δt\delta_t:

δt=Rt+1+γVϕ(St+1)−Vϕ(St)\delta_t = R_{t+1} + \gamma V_\phi(S_{t+1}) - V_\phi(S_t)

The scalar δt\delta_t serves as an unbiased estimate of the Advantage function A(St,At)=Q(St,At)−V(St)A(S_t, A_t) = Q(S_t, A_t) - V(S_t).

Joint Optimization Objectives

The network parameters are trained jointly using a multi-task composite loss function:

Ltotal(θ,ϕ)=Lactor(θ)+c1Lcritic(ϕ)−c2H(πθ(St))\mathcal{L}_{\text{total}}(\theta, \phi) = \mathcal{L}_{\text{actor}}(\theta) + c_1 \mathcal{L}_{\text{critic}}(\phi) - c_2 \mathcal{H}(\pi_\theta(S_t))

Where:

  • Actor Loss: Policy gradient weighted by the advantage: Lactor(θ)=−log⁡πθ(At∣St) δt\mathcal{L}_{\text{actor}}(\theta) = - \log \pi_\theta(A_t \mid S_t) \, \delta_t
  • Critic Loss: Mean squared error between predicted value and bootstrapped target: Lcritic(ϕ)=12(Rt+1+γVϕ(St+1)−Vϕ(St))2=12δt2\mathcal{L}_{\text{critic}}(\phi) = \frac{1}{2} \left( R_{t+1} + \gamma V_\phi(S_{t+1}) - V_\phi(S_t) \right)^2 = \frac{1}{2} \delta_t^2
  • Entropy Regularizer: H(πθ)=−∑aπ(a∣s)log⁡π(a∣s)\mathcal{H}(\pi_\theta) = -\sum_a \pi(a \mid s) \log \pi(a \mid s) encourages exploration and prevents premature policy collapse, weighted by coefficient c2c_2.

Coordination Schemes: A2C vs A3C

  • A3C (Asynchronous Advantage Actor-Critic): Mnih et al. (2016) introduced asynchronous parallel CPU workers. Multiple environment threads independently step through trajectories and push asynchronous, lock-free gradient updates (Hogwild!) to a centralized global network.
  • A2C (Advantage Actor-Critic): Synchronous batched implementation. A centralized coordinator collects fixed-step rollout batches (nn-steps) across vectorized parallel environments simultaneously, processing them as a single mini-batch on GPU hardware. A2C matches or exceeds A3C's empirical performance without asynchronous stale-gradient artifacts.

Worked Example

An agent observes state S0S_0. The shared network computes:

  • Critic evaluation: V(S0)=2.0V(S_0) = 2.0
  • Actor distribution over two discrete actions: π(a1∣S0)=0.70\pi(a_1 \mid S_0) = 0.70, π(a2∣S0)=0.30\pi(a_2 \mid S_0) = 0.30

The agent samples action A0=a1A_0 = a_1. The environment transitions to S1S_1 and awards R1=1.0R_1 = 1.0. The Critic evaluates next state: V(S1)=2.5V(S_1) = 2.5. Hyperparameters: discount factor γ=0.9\gamma = 0.9, learning rate α=0.1\alpha = 0.1.

  1. Calculate TD Advantage:

    δ0=R1+γV(S1)−V(S0)=1.0+0.9(2.5)−2.0=1.0+2.25−2.0=1.25\delta_0 = R_1 + \gamma V(S_1) - V(S_0) = 1.0 + 0.9(2.5) - 2.0 = 1.0 + 2.25 - 2.0 = 1.25

    Because δ0>0\delta_0 > 0, taking a1a_1 yielded higher returns than the Critic expected for state S0S_0.

  2. Critic Gradient Update:

    Target=R1+γV(S1)=3.25\text{Target} = R_1 + \gamma V(S_1) = 3.25 V(S0)←V(S0)+α×δ0=2.0+0.1×1.25=2.125V(S_0) \leftarrow V(S_0) + \alpha \times \delta_0 = 2.0 + 0.1 \times 1.25 = 2.125
  3. Actor Update Direction:

    • The log-probability of the chosen action a1a_1 is log⁡(0.70)≈−0.3567\log(0.70) \approx -0.3567.
    • The policy gradient ascends in proportion to ∇θlog⁡π(a1∣S0)×δ0=∇θlog⁡π(a1∣S0)×1.25\nabla_\theta \log \pi(a_1 \mid S_0) \times \delta_0 = \nabla_\theta \log \pi(a_1 \mid S_0) \times 1.25.
    • The probability of taking a1a_1 in S0S_0 will increase on the next training step.

Code

from typing import Tupleimport numpy as np
def compute_actor_critic_step(    v_current: float,    v_next: float,    prob_action: float,    reward: float,    gamma: float = 0.9,    alpha: float = 0.1,) -> Tuple[float, float, float]:    """Computes TD advantage and updated state value and policy log-prob delta."""    # 1-step TD advantage    td_target = reward + gamma * v_next    advantage = td_target - v_current
    # Critic update: move value toward td_target    updated_v = v_current + alpha * advantage
    # Actor policy gradient step: delta = alpha * (grad_log_pi * advantage)    # For a softmax logit update, the positive advantage encourages this action    logit_delta = alpha * advantage
    return advantage, updated_v, logit_delta
# Test case matching worked exampleadv, new_v, delta_logit = compute_actor_critic_step(    v_current=2.0,    v_next=2.5,    prob_action=0.7,    reward=1.0,    gamma=0.9,    alpha=0.1,)
print(f"Advantage (delta): {adv:.2f}")# -> Advantage (delta): 1.25
print(f"Updated V(S0): {new_v:.4f}")# -> Updated V(S0): 2.1250
print(f"Actor logit shift: +{delta_logit:.4f}")# -> Actor logit shift: +0.1250

Watch Out For

Asynchronous Worker Stale Gradients in A3C

In original A3C implementations, distributed asynchronous worker threads write parameter updates to a global server without locks. When workers operate at varying speeds or encounter complex simulation resets, a slow worker computes gradients with respect to parameter weights that are dozens of updates out of date. Applying these stale gradients destabilizes the policy and destroys the Critic's calibration.

In modern production systems, replace A3C with synchronous A2C or batched PPO. Synchronous execution coordinates all parallel environment runners into unified GPU tensor operations, guaranteeing zero parameter staleness while maximizing hardware throughput.

The Quick Version

  • Actor-Critic algorithms decouple decision-making (Actor πθ\pi_\theta) from performance evaluation (Critic VϕV_\phi).
  • The Critic computes a 1-step temporal difference error δt=Rt+1+γV(St+1)−V(St)\delta_t = R_{t+1} + \gamma V(S_{t+1}) - V(S_t) that acts as a low-variance baseline advantage.
  • Synchronous Advantage Actor-Critic (A2C) executes batched environments on GPUs, eliminating the stale-gradient instability of asynchronous A3C.