Actor-Critic Architecture
Actor-critic splits reinforcement learning into two cooperating networks: an Actor that selects actions and a Critic that grades those choices step-by-step using temporal difference errors.
Why Does This Exist?
Pure policy gradient methods like REINFORCE suffer from excessive variance because they evaluate actions using complete trajectory returns . In long-horizon tasks, small variations in environment randomness thousands of steps downstream corrupt the gradient of an action taken at step zero. Conversely, pure value-based methods like Q-learning bootstrap efficiently on single-step transitions but cannot natively parameterize continuous action spaces or learn stochastic policies.
The Actor-Critic architecture unites both families into a single hybrid system. An Actor parameterizes the policy , while a Critic parameterizes a value function .
Instead of waiting for an entire episode to finish, the Actor updates after every transition (or over small -step rollout windows) using temporal difference (TD) errors provided by the Critic. By replacing the noisy Monte Carlo return with a 1-step bootstrapped estimate , Actor-Critic methods drastically slash variance while retaining the ability to optimize arbitrary continuous or discrete policies.
Think of It Like This
A stage actor rehearsing with a director giving instant notes
Imagine a theater performer (the Actor) rehearsing lines for an upcoming play.
Under a pure Monte Carlo approach (REINFORCE), the performer would deliver all five acts of the play from start to finish, wait until opening night to see the final newspaper review, and only then try to recall which specific facial expressions in Act 1 contributed to the critic's overall verdict. Learning this way is excruciatingly slow and full of second-guessing.
Under an Actor-Critic setup, a professional director (the Critic) sits in the front row during rehearsals. The performer delivers a single line (takes action ). Immediately, the director interrupts: "That inflection was sharper than usual, keep that" (positive TD error), or "That tone fell flat compared to what this scene requires" (negative TD error).
The performer immediately adjusts their line delivery (updates policy parameters ) based on the director's localized feedback. Meanwhile, the director continually refines their own mental standard of how much applause each scene is worth (updates value parameters ).
How It Actually Works
Advantage Estimation and Dual Objective Updates
In modern deep Actor-Critic architectures, the Actor and Critic typically share a common convolutional or transformer representation backbone , branching into two output heads:
- Actor Head: Outputs action distribution parameters (e.g., categorical logits or Gaussian mean and standard deviation).
- Critic Head: Outputs a scalar state-value baseline .
When the agent executes action , receives reward , and transitions to , the Critic computes the 1-step Temporal Difference Advantage :
The scalar serves as an unbiased estimate of the Advantage function .
Joint Optimization Objectives
The network parameters are trained jointly using a multi-task composite loss function:
Where:
- Actor Loss: Policy gradient weighted by the advantage:
- Critic Loss: Mean squared error between predicted value and bootstrapped target:
- Entropy Regularizer: encourages exploration and prevents premature policy collapse, weighted by coefficient .
Coordination Schemes: A2C vs A3C
- A3C (Asynchronous Advantage Actor-Critic): Mnih et al. (2016) introduced asynchronous parallel CPU workers. Multiple environment threads independently step through trajectories and push asynchronous, lock-free gradient updates (Hogwild!) to a centralized global network.
- A2C (Advantage Actor-Critic): Synchronous batched implementation. A centralized coordinator collects fixed-step rollout batches (-steps) across vectorized parallel environments simultaneously, processing them as a single mini-batch on GPU hardware. A2C matches or exceeds A3C's empirical performance without asynchronous stale-gradient artifacts.
Worked Example
An agent observes state . The shared network computes:
- Critic evaluation:
- Actor distribution over two discrete actions: ,
The agent samples action . The environment transitions to and awards . The Critic evaluates next state: . Hyperparameters: discount factor , learning rate .
-
Calculate TD Advantage:
Because , taking yielded higher returns than the Critic expected for state .
-
Critic Gradient Update:
-
Actor Update Direction:
- The log-probability of the chosen action is .
- The policy gradient ascends in proportion to .
- The probability of taking in will increase on the next training step.
Code
from typing import Tupleimport numpy as np
def compute_actor_critic_step( v_current: float, v_next: float, prob_action: float, reward: float, gamma: float = 0.9, alpha: float = 0.1,) -> Tuple[float, float, float]: """Computes TD advantage and updated state value and policy log-prob delta.""" # 1-step TD advantage td_target = reward + gamma * v_next advantage = td_target - v_current
# Critic update: move value toward td_target updated_v = v_current + alpha * advantage
# Actor policy gradient step: delta = alpha * (grad_log_pi * advantage) # For a softmax logit update, the positive advantage encourages this action logit_delta = alpha * advantage
return advantage, updated_v, logit_delta
# Test case matching worked exampleadv, new_v, delta_logit = compute_actor_critic_step( v_current=2.0, v_next=2.5, prob_action=0.7, reward=1.0, gamma=0.9, alpha=0.1,)
print(f"Advantage (delta): {adv:.2f}")# -> Advantage (delta): 1.25
print(f"Updated V(S0): {new_v:.4f}")# -> Updated V(S0): 2.1250
print(f"Actor logit shift: +{delta_logit:.4f}")# -> Actor logit shift: +0.1250Watch Out For
Asynchronous Worker Stale Gradients in A3C
In original A3C implementations, distributed asynchronous worker threads write parameter updates to a global server without locks. When workers operate at varying speeds or encounter complex simulation resets, a slow worker computes gradients with respect to parameter weights that are dozens of updates out of date. Applying these stale gradients destabilizes the policy and destroys the Critic's calibration.
In modern production systems, replace A3C with synchronous A2C or batched PPO. Synchronous execution coordinates all parallel environment runners into unified GPU tensor operations, guaranteeing zero parameter staleness while maximizing hardware throughput.
The Quick Version
- Actor-Critic algorithms decouple decision-making (Actor ) from performance evaluation (Critic ).
- The Critic computes a 1-step temporal difference error that acts as a low-variance baseline advantage.
- Synchronous Advantage Actor-Critic (A2C) executes batched environments on GPUs, eliminating the stale-gradient instability of asynchronous A3C.