Continuous Action Spaces in RL
Instead of picking from a rigid list of discrete options, continuous action policies output smooth, real-valued control signals like motor torques, steering angles, and throttle pressures.
Why Does This Exist?
In classical reinforcement learning, algorithms like Q-learning and Deep Q-Networks (DQN) operate on discrete action spaces . To select the optimal action, the agent evaluates the state-action value for every available action and extracts the maximum:
However, real-world physical systems—such as robotic manipulators, autonomous quadcopters, industrial process valves, and self-driving vehicles—operate in continuous action domains , where control signals represent physical quantities like joint torques, motor voltages, and steering angles.
Attempting to apply discrete RL to continuous domains by discretizing each action dimension triggers two critical failure modes:
- The Curse of Dimensionality: If an -degree-of-freedom robotic arm discretizes each joint into discrete bins, the total action space explodes exponentially to . For a standard 7-joint manipulator with a modest bins per joint, the agent must evaluate distinct Q-values at every single control step. Computing over tens of millions of forward passes becomes computationally intractable.
- Loss of Physical Smoothness and Actuator Damage: Discretization replaces smooth physical dynamics with coarse, stepped commands. A simulated robot trying to balance with discrete bang-bang control oscillates violently around equilibrium, inducing high-frequency chatter that would strip real physical gears and overheat motors.
Continuous action space methods resolve these limitations by parameterizing policies directly as continuous probability density functions (stochastic continuous control) or smooth mapping functions (deterministic continuous control). Instead of searching over discrete actions, neural networks directly output the continuous control vector in a single forward pass.
Think of It Like This
Steering a Yacht: Smooth Helm vs Discrete Gamepad Buttons
Imagine steering a 40-foot racing yacht through a narrow, windswept harbor channel.
If your helm control was restricted to a digital arcade gamepad (a discrete action space), you would only have three buttons: Hard Port (Left), Dead Center, and Hard Starboard (Right). To maintain a heading of , you would have to rapidly oscillate between jamming the rudder hard left and hard right. The yacht would violently lurch from side to side, shedding hydrodynamic momentum, straining the steering cables, and risking capsize.
A real ship uses a continuous ship's wheel (a continuous action space). The helmsman can rotate the helm smoothly to any precise real-valued angle—, , or —and hold it effortlessly. The rudder creates smooth laminar water flow, effortlessly tracking the true target heading.
Where the analogy stops: A ship's wheel is a single 1D physical axle, whereas continuous robot controllers manipulate high-dimensional multidimensional vectors simultaneously (e.g., continuous joint torques on a humanoid walking robot), requiring multidimensional probability density modeling and cross-joint coordination.
How It Actually Works
Parameterizing Stochastic Continuous Policies
The standard framework for stochastic continuous control parameterizes the policy as a multivariate diagonal Gaussian distribution:
The probability density of selecting action in state is the product of independent 1D Gaussians:
Taking the natural logarithm yields the log-likelihood:
To implement this with deep neural networks:
- The network features split into two heads: a mean head and a log standard deviation head .
- Predicting rather than raw guarantees that strictly holds without numerical instability.
- To prevent policy entropy from exploding or collapsing, is clamped to a stable numerical bracket (typically ).
The Two Policy Gradient Paradigms
Continuous action policies are trained using one of two fundamental gradient mechanisms:
1. The Score Function / Likelihood Ratio Method (REINFORCE, PPO, TRPO)
The policy gradient theorem establishes that the gradient of expected return is:
For a diagonal Gaussian, the analytical score function decomposes cleanly:
- If an action yields positive advantage , the mean shifts toward .
- If the sampled action is farther from than one standard deviation (), the variance expands to encourage further exploration; otherwise, it contracts.
2. The Pathwise Derivative / Reparameterization Trick (SAC)
When an action-value critic is available and differentiable with respect to , the score function exhibits unnecessarily high variance. Instead, we use the reparameterization trick (Kingma & Welling, 2013).
Instead of sampling directly from , we express the action as a deterministic, differentiable transformation of an independent standard normal noise variable :
Now, the expectation over the stochastic policy can be written over the static noise distribution: Gradients flow directly through the action sample into the policy parameters via the chain rule, slashing gradient variance by orders of magnitude.
Tanh Squashing and the Mandatory Jacobian Correction
Physical actuators have hard operational limits (e.g., maximum torque ). Standard Gaussian distributions have infinite support . Clipping the action destroys gradients outside the boundaries because whenever .
Modern continuous RL algorithms (such as Soft Actor-Critic / SAC) solve this by applying an invertible, smooth non-linear squashing function:
However, non-linear squashing compresses the volume of probability space near the boundaries. To compute the true log probability required for entropy regularization, we must apply the multidimensional change-of-variables theorem:
Because operates coordinate-wise, the Jacobian matrix is diagonal:
Taking the logarithm yields the mandatory Jacobian log-determinant correction: where provides numerical stabilization against division by zero when .
Worked numerical example
Let us calculate a 1D continuous policy evaluation step with Gaussian parameterization and squashing:
- Mean:
- Standard deviation: (variance )
- Action bound:
- Evaluated latent sample:
Step 1: Base Gaussian density and log-probability
Compute log density:
- Normalization term:
- Scale term:
- Mahalanobis exponent:
Sum the terms: Raw Gaussian density:
Step 2: Score function gradients
- Gradient with respect to mean :
- Gradient with respect to :
(Because the deviation is exactly equal to one standard deviation , the score for is zero—the distribution's current variance perfectly matches the sampled point).
Step 3: Tanh squashing and Jacobian log-prob correction
Squash the latent variable into physical action space:
Compute the local derivative of :
Compute the log-determinant of the Jacobian:
Apply the change-of-variables correction:
Squashed probability density:
Notice that the probability density is much higher than the raw Gaussian density . Because heavily contracts high-magnitude inputs into a tiny interval near , the local probability density is compressed and concentrated.
Code
The following self-contained Python script implements a Squashed Gaussian Policy with pathwise reparameterization, numerical clamping, and change-of-variables Jacobian correction:
"""Continuous Gaussian Policy with Tanh Squashing and Jacobian Correction.
Implements the standard policy distribution used in modern continuous actor-criticarchitectures like Soft Actor-Critic (SAC)."""
from typing import Dict, Tupleimport numpy as np
class SquashedGaussianPolicy: """Continuous policy outputting bounded actions in [-action_scale, action_scale]."""
LOG_STD_MIN: float = -20.0 LOG_STD_MAX: float = 2.0 EPSILON: float = 1e-6
def __init__( self, state_dim: int = 3, action_dim: int = 2, action_scale: float = 2.0, seed: int = 42 ) -> None: np.random.seed(seed) self.state_dim: int = state_dim self.action_dim: int = action_dim self.action_scale: float = action_scale
# Shared feature trunk self.W1: np.ndarray = np.random.randn(state_dim, 32) * 0.1 self.b1: np.ndarray = np.zeros(32)
# Dual output heads: Mean and Log-Std self.W_mu: np.ndarray = np.random.randn(32, action_dim) * 0.1 self.b_mu: np.ndarray = np.zeros(action_dim)
self.W_log_std: np.ndarray = np.random.randn(32, action_dim) * 0.1 self.b_log_std: np.ndarray = np.zeros(action_dim)
def forward(self, state: np.ndarray) -> Tuple[np.ndarray, np.ndarray]: """Compute Gaussian mean and clamped log standard deviation.""" h = np.maximum(0, state @ self.W1 + self.b1) mu = h @ self.W_mu + self.b_mu log_std = h @ self.W_log_std + self.b_log_std
# Clamp log standard deviation to avoid numerical instability log_std = np.clip(log_std, self.LOG_STD_MIN, self.LOG_STD_MAX) return mu, log_std
def sample( self, state: np.ndarray, deterministic: bool = False ) -> Tuple[np.ndarray, float, np.ndarray]: """Sample bounded action via reparameterization trick and Jacobian correction. Returns: action: Bounded continuous action in [-action_scale, action_scale]. corrected_log_prob: Exact change-of-variables log probability. raw_u: Unsquashed latent Gaussian variable. """ mu, log_std = self.forward(state) std = np.exp(log_std)
if deterministic: u = mu else: # Reparameterization trick: u = mu + sigma * epsilon epsilon = np.random.randn(*mu.shape) u = mu + std * epsilon
# 1. Raw Gaussian log probability: log p(u | s) var = std**2 gaussian_log_prob = -0.5 * np.sum( ((u - mu)**2) / var + 2.0 * log_std + np.log(2.0 * np.pi) )
# 2. Tanh squashing and scaling a_unit = np.tanh(u) action = a_unit * self.action_scale
# 3. Jacobian log-determinant correction: # da / du = action_scale * (1 - tanh^2(u)) # log |det J| = sum( log(action_scale) + log(1 - tanh^2(u) + eps) ) log_det_jacobian = np.sum( np.log(self.action_scale) + np.log(1.0 - a_unit**2 + self.EPSILON) )
# 4. Corrected log probability corrected_log_prob = float(gaussian_log_prob - log_det_jacobian)
return action, corrected_log_prob, u
def run_continuous_demo() -> Dict[str, float]: """Execute end-to-end verification of continuous Gaussian policy.""" policy = SquashedGaussianPolicy(state_dim=3, action_dim=2, action_scale=2.0, seed=42) state = np.array([1.0, -0.5, 0.2])
stochastic_action, log_prob, u = policy.sample(state, deterministic=False) deterministic_action, _, _ = policy.sample(state, deterministic=True)
return { "action_0": float(stochastic_action[0]), "action_1": float(stochastic_action[1]), "log_prob": float(log_prob), "det_0": float(deterministic_action[0]), "det_1": float(deterministic_action[1]), }
if __name__ == "__main__": res = run_continuous_demo() print("Continuous Policy Outputs:") print(f"Stochastic Action: [{res['action_0']:.4f}, {res['action_1']:.4f}]") print(f"Corrected Log Prob: {res['log_prob']:.4f}") print(f"Deterministic Mode: [{res['det_0']:.4f}, {res['det_1']:.4f}]")
# Automated assertions assert -2.0 <= res["action_0"] <= 2.0, "Action 0 must respect physical bounds [-2, 2]." assert -2.0 <= res["action_1"] <= 2.0, "Action 1 must respect physical bounds [-2, 2]." assert np.isfinite(res["log_prob"]), "Log probability must be a finite numerical value." assert -2.0 <= res["det_0"] <= 2.0, "Deterministic mode must respect physical bounds." print("Verification passed: Continuous policy correctly enforces bounds and Jacobian corrections.")Expected Output
Continuous Policy Outputs:Stochastic Action: [-1.0112, 1.6313]Corrected Log Prob: -2.6451Deterministic Mode: [-0.1327, -0.0439]Verification passed: Continuous policy correctly enforces bounds and Jacobian corrections.Watch Out For
The Missing Jacobian Trap and Premature Entropy Collapse
The Trap: Practitioners frequently encounter two silent failure modes in continuous policy gradients:
- The Missing Jacobian Correction: In algorithms that use entropy regularization (such as Soft Actor-Critic), computing policy entropy using the raw Gaussian probability instead of the squashed probability completely corrupts the entropy objective. Because compresses actions near , failing to subtract falsely incentives the policy to push its pre-squashed distribution to extreme magnitudes (), causing gradients to saturate and training to freeze.
- Premature Entropy Collapse: If the log standard deviation head is unconstrained, early training noise can cause , forcing . The policy collapses into a deterministic Dirac delta distribution before it has sufficiently explored the state space, permanently trapping the agent in suboptimal local minima.
The Fix:
- Always apply the change-of-variables correction: .
- Strictly bound the standard deviation head by clipping . This ensures the policy maintains a nonzero minimum exploration variance () while preventing explosive variance.
The Quick Version
- Solves Combinatorial Discretization: Directly emits smooth real-valued control vectors (e.g., torques and voltages), avoiding the exponential action explosion of discrete grids.
- Gaussian Parameterization: Neural networks output the distribution mean and clamped log standard deviation , guaranteeing valid positive variance.
- Reparameterization Trick: Decomposes sampling into , allowing pathwise policy gradients to backpropagate directly through actions with minimal variance.
- Tanh Squashing & Jacobian Correction: Restricts actions to physical motor bounds ; requires subtracting the log-determinant to maintain mathematically exact probability densities.