Soft Actor-Critic (SAC)
SAC balances reward maximization with policy entropy, using twin critics and a reparameterized Gaussian actor for stable, sample-efficient continuous control.
Why Does This Exist?
In continuous control reinforcement learning, deterministic off-policy algorithms like Deep Deterministic Policy Gradient (DDPG) often suffer from hyperparameter sensitivity and severe value overestimation. Because deterministic policies converge prematurely to local optima, small errors in action-value estimation compound, driving the policy toward sub-optimal, brittle behaviors.
Conversely, on-policy algorithms such as Proximal Policy Optimization (PPO) provide stability through stochastic policies, but discard historical transitions after every update batch, requiring millions of interactions with physical or simulated environments.
Soft Actor-Critic (SAC), introduced by Haarnoja et al. (2018), bridges this divide. By framing policy optimization under the Maximum Entropy RL framework, SAC instructs the agent to maximize both cumulative task reward and policy entropy. This entropy incentive encourages broad exploration, prevents policy collapse onto near-deterministic narrow trajectories, and yields robust policies capable of adapting to environmental perturbations while retaining full off-policy sample efficiency.
Think of It Like This
A Gymnast Balancing on the High Bar
Picture a gymnast mastering an intricate release-and-regrasp routine on the high bar.
A rigid coach (standard deterministic policy gradient) demands an exact millimeter-by-millimeter trajectory. The gymnast memorizes a single, fragile sequence. During competition, an imperceptible draft or a slick spot on the bar causes a microscopic twitch. Because the gymnast only knows one exact sequence, the unexpected tremor leads directly to a fall.
A modern acrobatic trainer (Soft Actor-Critic) scores routines on two simultaneous criteria: technical difficulty (task reward) and stylistic fluidity with adaptive recovery (entropy bonus). Rather than drilling a single rigid line, the gymnast is incentivized to master an entire family of viable trajectories around the bar.
When a sudden gust or grip slippage occurs, the gymnast does not panic or collapse; their repertoire includes multiple fluid corrections that keep them on the bar. Over time, as difficulty peaks, the trainer gently modulates the focus from exploration to pure execution (automatic temperature tuning), ensuring top competition scores.
Where the analogy stops: A human gymnast relies on biological proprioception and muscle memory. SAC represents policy flexibility mathematically via a squashed Gaussian distribution whose variance and covariance are parameterized by neural networks and dynamically regularized by a scalar entropy temperature .
How It Actually Works
Maximum Entropy Objective and the Tri-Component Architecture
Standard reinforcement learning maximizes expected return . Soft Actor-Critic generalizes this to the maximum entropy objective:
where is the Shannon entropy of the action distribution, and is the entropy temperature determining the relative importance of the entropy bonus against reward maximization.
SAC implements this optimization through three interacting mechanisms:
-
Clipped Double Q-Learning with Soft Bellman Backups: To prevent Q-function overestimation bias, SAC maintains two independent critic networks, and , alongside slowly moving target critics and . The soft state-value target takes the minimum between the twin critics and penalizes high certainty (low entropy):
Both critics are trained by minimizing mean squared Bellman error:
-
Reparameterized Squashed Gaussian Policy: Continuous actions are bounded to a valid physical range using the hyperbolic tangent function . To allow backpropagation directly through policy sampling, actions are sampled via the reparameterization trick:
The policy parameters are optimized to maximize the expected soft Q-value:
Reparameterization Trick, Tanh Squashing, and Jacobian Correction
Because is a nonlinear change of variables from the latent Gaussian variable to the bounded action , calculating requires applying the change-of-variables formula:
The derivative of the hyperbolic tangent is . Because the transformation is element-wise, the Jacobian matrix is diagonal, and its determinant is the product of diagonal entries. Taking the logarithm gives:
where and is a tiny stabilizer (e.g., ) to avoid taking the logarithm of zero when actions reach the boundary .
Automatic Entropy Temperature Tuning
Early iterations of SAC treated as a fixed hyperparameter. However, the optimal entropy varies significantly across environments and training phases: early exploration requires high entropy, while fine motor control requires lower entropy.
SAC resolves this by framing entropy as an inequality constraint: , where is the heuristic target entropy. This yields the dual optimization problem for :
Optimizing with gradient descent:
If policy entropy exceeds the target , automatically decreases. If the policy becomes overly deterministic, increases to stimulate exploration.
Worked numerical example
Let us trace a single transition update with concrete numbers:
Suppose an agent in 1D continuous action space (, target entropy ) transitions to next state . The policy produces mean and standard deviation . Suppose sampled noise .
-
Sample Latent and Squashed Action:
-
Evaluate Jacobian and Log-Probabilities:
- Tanh Jacobian factor:
- Latent Gaussian log-probability:
- Squashed corrected log-probability:
-
Compute Soft Bellman Target: Let current temperature , reward , discount . The twin target critics evaluate the pair as:
Soft state value with entropy deduction:
Bellman target :
-
Update Temperature : Compare current negative log-prob with target entropy :
Entropy difference: . Because , the policy has more entropy than the target constraint demands. The gradient on is positive, triggering a decrease in to focus policy execution.
Code
import numpy as np
class SoftActorCriticStep: """Implements core mathematics of Soft Actor-Critic (SAC) updates:
Reparameterized squashed Gaussian sampling, Jacobian correction, clipped double Q soft Bellman targets, and automatic temperature tuning. """
def __init__( self, action_dim: int = 1, gamma: float = 0.99, initial_alpha: float = 0.2, eps_stabilizer: float = 1e-6, ) -> None: self.action_dim = action_dim self.gamma = gamma self.log_alpha = float(np.log(initial_alpha)) self.target_entropy = -float(action_dim) self.eps = eps_stabilizer
@property def alpha(self) -> float: return float(np.exp(self.log_alpha))
def sample_action( self, mu: np.ndarray, sigma: np.ndarray, noise: np.ndarray | None = None, ) -> tuple[np.ndarray, np.ndarray, float]: """Samples action via reparameterization trick: u = mu + sigma * eps, a = tanh(u).
Applies Jacobian change-of-variables correction to calculate log pi(a|s). """ if noise is None: noise = np.random.randn(*mu.shape)
# Reparameterization trick: differentiable latent Gaussian sample u = mu + sigma * noise a = np.tanh(u)
# Gaussian density in unconstrained latent space u var = sigma**2 log_p_u = -0.5 * np.sum(np.log(2.0 * np.pi * var) + ((u - mu) ** 2) / var)
# Jacobian change-of-variables correction for tanh squashing log_jacobian = float(np.sum(np.log(1.0 - a**2 + self.eps))) log_pi = float(log_p_u - log_jacobian)
return u, a, log_pi
def compute_soft_target( self, reward: float, q1_target: float, q2_target: float, next_log_pi: float, ) -> float: """Computes soft Bellman value target with clipped double Q and entropy bonus.""" min_q = min(q1_target, q2_target) soft_value = min_q - self.alpha * next_log_pi y = reward + self.gamma * soft_value return float(y)
def compute_temperature_update( self, log_pi: float, lr_alpha: float = 0.05, ) -> tuple[float, float, float]: """Computes automatic entropy temperature loss, gradient, and updated alpha.""" entropy_diff = -log_pi - self.target_entropy alpha_loss = -self.log_alpha * (log_pi + self.target_entropy) grad_log_alpha = entropy_diff
# Gradient descent step on log(alpha) new_log_alpha = self.log_alpha - lr_alpha * grad_log_alpha new_alpha = float(np.exp(new_log_alpha))
return alpha_loss, grad_log_alpha, new_alpha
if __name__ == "__main__": np.set_printoptions(precision=4, suppress=True)
# Initialize SAC step for 1D continuous action space sac = SoftActorCriticStep(action_dim=1, gamma=0.99, initial_alpha=0.2)
# Transition parameters matching the worked numerical example mu = np.array([1.0], dtype=np.float64) sigma = np.array([0.5], dtype=np.float64) noise = np.array([0.0], dtype=np.float64)
u, a, log_pi = sac.sample_action(mu, sigma, noise=noise)
q1_next, q2_next = 5.0, 4.2 reward = 1.0 soft_target = sac.compute_soft_target(reward, q1_next, q2_next, log_pi)
alpha_loss, grad_log_alpha, updated_alpha = sac.compute_temperature_update( log_pi, lr_alpha=0.05 )
print(f"Latent sample u: {u[0]:.4f}") # -> Latent sample u: 1.0000
print(f"Squashed action a: {a[0]:.4f}") # -> Squashed action a: 0.7616
print(f"Corrected log pi(a|s): {log_pi:.4f}") # -> Corrected log pi(a|s): 0.6418
print(f"Soft Bellman target y: {soft_target:.4f}") # -> Soft Bellman target y: 5.0309
print(f"Temperature alpha: {sac.alpha:.4f}") # -> Temperature alpha: 0.2000
print(f"Target entropy H_bar: {sac.target_entropy:.4f}") # -> Target entropy H_bar: -1.0000
print(f"Log-alpha gradient: {grad_log_alpha:.4f}") # -> Log-alpha gradient: 0.3582
print(f"Updated alpha: {updated_alpha:.4f}") # -> Updated alpha: 0.1964
# Assert correctness assert np.isclose(a[0], 0.7616, atol=1e-3) assert np.isclose(soft_target, 5.0309, atol=1e-3) assert updated_alpha < sac.alpha # Alpha reduced as entropy > targetWatch Out For
Omitting the Tanh Jacobian Change-of-Variables Correction
A common bug in custom SAC implementations is computing using the raw Gaussian probability density function without subtracting the Jacobian determinant penalty .
Because compresses into , probability mass piles up heavily near the action boundaries . Without the Jacobian correction, the policy's calculated entropy is vastly overestimated near the boundaries, leading the automatic temperature tuner to drive prematurely or causing severe gradient explosion during backpropagation. Always apply the change-of-variables formula with numerical stabilization .
The Quick Version
- Maximum Entropy Framework: SAC optimizes for both cumulative rewards and policy entropy, encouraging wide exploration and preventing convergence to brittle, deterministic policies.
- Clipped Double Q-Learning: Uses the minimum prediction of two independent target critics () in soft Bellman backups to suppress value overestimation.
- Reparameterized Tanh Actor: Enables end-to-end backpropagation through stochastic action choices while bounding continuous actions to with an analytical Jacobian correction.
- Automatic Temperature Tuning: Continuously modulates the entropy weight to satisfy an entropy constraint , balancing exploration early and exploitation late.