Skip to content
AI360Xpert
Beta

Soft Actor-Critic (SAC)

SAC balances reward maximization with policy entropy, using twin critics and a reparameterized Gaussian actor for stable, sample-efficient continuous control.

Soft Actor-Critic couples clipped twin critics with a reparameterized squashed Gaussian actor and dynamic entropy temperature tuning for stable continuous control.
Soft Actor-Critic couples clipped twin critics with a reparameterized squashed Gaussian actor and dynamic entropy temperature tuning for stable continuous control.

Why Does This Exist?

In continuous control reinforcement learning, deterministic off-policy algorithms like Deep Deterministic Policy Gradient (DDPG) often suffer from hyperparameter sensitivity and severe value overestimation. Because deterministic policies converge prematurely to local optima, small errors in action-value estimation compound, driving the policy toward sub-optimal, brittle behaviors.

Conversely, on-policy algorithms such as Proximal Policy Optimization (PPO) provide stability through stochastic policies, but discard historical transitions after every update batch, requiring millions of interactions with physical or simulated environments.

Soft Actor-Critic (SAC), introduced by Haarnoja et al. (2018), bridges this divide. By framing policy optimization under the Maximum Entropy RL framework, SAC instructs the agent to maximize both cumulative task reward and policy entropy. This entropy incentive encourages broad exploration, prevents policy collapse onto near-deterministic narrow trajectories, and yields robust policies capable of adapting to environmental perturbations while retaining full off-policy sample efficiency.

Think of It Like This

A Gymnast Balancing on the High Bar

Picture a gymnast mastering an intricate release-and-regrasp routine on the high bar.

A rigid coach (standard deterministic policy gradient) demands an exact millimeter-by-millimeter trajectory. The gymnast memorizes a single, fragile sequence. During competition, an imperceptible draft or a slick spot on the bar causes a microscopic twitch. Because the gymnast only knows one exact sequence, the unexpected tremor leads directly to a fall.

A modern acrobatic trainer (Soft Actor-Critic) scores routines on two simultaneous criteria: technical difficulty (task reward) and stylistic fluidity with adaptive recovery (entropy bonus). Rather than drilling a single rigid line, the gymnast is incentivized to master an entire family of viable trajectories around the bar.

When a sudden gust or grip slippage occurs, the gymnast does not panic or collapse; their repertoire includes multiple fluid corrections that keep them on the bar. Over time, as difficulty peaks, the trainer gently modulates the focus from exploration to pure execution (automatic temperature tuning), ensuring top competition scores.

Where the analogy stops: A human gymnast relies on biological proprioception and muscle memory. SAC represents policy flexibility mathematically via a squashed Gaussian distribution whose variance and covariance are parameterized by neural networks and dynamically regularized by a scalar entropy temperature α\alpha.

How It Actually Works

Maximum Entropy Objective and the Tri-Component Architecture

Standard reinforcement learning maximizes expected return ∑tE[r(st,at)]\sum_t \mathbb{E}[r(s_t, a_t)]. Soft Actor-Critic generalizes this to the maximum entropy objective:

J(π)=∑t=0TE(st,at)∼ρπ[r(st,at)+αH(π(⋅∣st))]J(\pi) = \sum_{t=0}^T \mathbb{E}_{(s_t, a_t) \sim \rho_\pi} \left[ r(s_t, a_t) + \alpha \mathcal{H}(\pi(\cdot | s_t)) \right]

where H(π(⋅∣st))=Ea∼π[−log⁡π(a∣st)]\mathcal{H}(\pi(\cdot | s_t)) = \mathbb{E}_{a \sim \pi}[-\log \pi(a | s_t)] is the Shannon entropy of the action distribution, and α\alpha is the entropy temperature determining the relative importance of the entropy bonus against reward maximization.

SAC implements this optimization through three interacting mechanisms:

  1. Clipped Double Q-Learning with Soft Bellman Backups: To prevent Q-function overestimation bias, SAC maintains two independent critic networks, Qθ1Q_{\theta_1} and Qθ2Q_{\theta_2}, alongside slowly moving target critics Qθˉ1Q_{\bar{\theta}_1} and Qθˉ2Q_{\bar{\theta}_2}. The soft state-value target takes the minimum between the twin critics and penalizes high certainty (low entropy):

    y=r+γ(min⁡i=1,2Qθˉi(s′,a~′)−αlog⁡πϕ(a~′∣s′)),a~′∼πϕ(⋅∣s′)y = r + \gamma \left( \min_{i=1,2} Q_{\bar{\theta}_i}(s', \tilde{a}') - \alpha \log \pi_\phi(\tilde{a}' | s') \right), \quad \tilde{a}' \sim \pi_\phi(\cdot | s')

    Both critics are trained by minimizing mean squared Bellman error:

    L(θi)=E(s,a,r,s′)∼D[(Qθi(s,a)−y)2]L(\theta_i) = \mathbb{E}_{(s, a, r, s') \sim \mathcal{D}} \left[ \left( Q_{\theta_i}(s, a) - y \right)^2 \right]
  2. Reparameterized Squashed Gaussian Policy: Continuous actions are bounded to a valid physical range [−1,1][-1, 1] using the hyperbolic tangent function tanh⁡\tanh. To allow backpropagation directly through policy sampling, actions are sampled via the reparameterization trick:

    u=μϕ(s)+σϕ(s)⊙ϵ,ϵ∼N(0,I)u = \mu_\phi(s) + \sigma_\phi(s) \odot \epsilon, \quad \epsilon \sim \mathcal{N}(0, I) a=tanh⁡(u)a = \tanh(u)

    The policy parameters ϕ\phi are optimized to maximize the expected soft Q-value:

    L(ϕ)=Es∼D,ϵ∼N[αlog⁡πϕ(a∣s)−min⁡i=1,2Qθi(s,a)]L(\phi) = \mathbb{E}_{s \sim \mathcal{D}, \epsilon \sim \mathcal{N}} \left[ \alpha \log \pi_\phi(a | s) - \min_{i=1,2} Q_{\theta_i}(s, a) \right]

Reparameterization Trick, Tanh Squashing, and Jacobian Correction

Because a=tanh⁡(u)a = \tanh(u) is a nonlinear change of variables from the latent Gaussian variable uu to the bounded action aa, calculating log⁡π(a∣s)\log \pi(a|s) requires applying the change-of-variables formula:

π(a∣s)=p(u∣s)∣det⁡(dadu)∣−1\pi(a | s) = p(u | s) \left| \det \left( \frac{da}{du} \right) \right|^{-1}

The derivative of the hyperbolic tangent is daidui=1−tanh⁡2(ui)=1−ai2\frac{da_i}{du_i} = 1 - \tanh^2(u_i) = 1 - a_i^2. Because the transformation is element-wise, the Jacobian matrix is diagonal, and its determinant is the product of diagonal entries. Taking the logarithm gives:

log⁡π(a∣s)=log⁡p(u∣s)−∑i=1dlog⁡(1−tanh⁡2(ui)+δ)\log \pi(a | s) = \log p(u | s) - \sum_{i=1}^{d} \log \left( 1 - \tanh^2(u_i) + \delta \right)

where p(u∣s)=N(μϕ(s),σϕ(s)2)p(u | s) = \mathcal{N}(\mu_\phi(s), \sigma_\phi(s)^2) and δ>0\delta > 0 is a tiny stabilizer (e.g., 10−610^{-6}) to avoid taking the logarithm of zero when actions reach the boundary [−1,1][-1, 1].

Automatic Entropy Temperature Tuning

Early iterations of SAC treated α\alpha as a fixed hyperparameter. However, the optimal entropy varies significantly across environments and training phases: early exploration requires high entropy, while fine motor control requires lower entropy.

SAC resolves this by framing entropy as an inequality constraint: E[−log⁡π(a∣s)]≥Hˉ\mathbb{E}[-\log \pi(a|s)] \ge \bar{\mathcal{H}}, where Hˉ=−dim⁡(A)\bar{\mathcal{H}} = -\dim(\mathcal{A}) is the heuristic target entropy. This yields the dual optimization problem for α\alpha:

J(α)=Es∼D,a∼π[−αlog⁡π(a∣s)−αHˉ]=E[−α(log⁡π(a∣s)+Hˉ)]J(\alpha) = \mathbb{E}_{s \sim \mathcal{D}, a \sim \pi} \left[ -\alpha \log \pi(a | s) - \alpha \bar{\mathcal{H}} \right] = \mathbb{E} \left[ -\alpha (\log \pi(a | s) + \bar{\mathcal{H}}) \right]

Optimizing log⁡α\log \alpha with gradient descent:

∇log⁡αL(α)=−α(log⁡π(a∣s)+Hˉ)\nabla_{\log \alpha} L(\alpha) = -\alpha (\log \pi(a | s) + \bar{\mathcal{H}})

If policy entropy −log⁡π(a∣s)-\log \pi(a|s) exceeds the target Hˉ\bar{\mathcal{H}}, α\alpha automatically decreases. If the policy becomes overly deterministic, α\alpha increases to stimulate exploration.

Worked numerical example

Let us trace a single transition update with concrete numbers:

Suppose an agent in 1D continuous action space (dim⁡(A)=1\dim(\mathcal{A}) = 1, target entropy Hˉ=−1.0\bar{\mathcal{H}} = -1.0) transitions to next state s′s'. The policy produces mean μ=1.0\mu = 1.0 and standard deviation σ=0.5\sigma = 0.5. Suppose sampled noise ϵ=0.0\epsilon = 0.0.

  1. Sample Latent and Squashed Action:

    u=μ+σϵ=1.0+(0.5)(0.0)=1.0u = \mu + \sigma \epsilon = 1.0 + (0.5)(0.0) = 1.0 a′=tanh⁡(1.0)≈0.7616a' = \tanh(1.0) \approx 0.7616
  2. Evaluate Jacobian and Log-Probabilities:

    • Tanh Jacobian factor: 1−tanh⁡2(1.0)=1−(0.7616)2=1−0.5800=0.42001 - \tanh^2(1.0) = 1 - (0.7616)^2 = 1 - 0.5800 = 0.4200 ln⁡(0.4200)≈−0.8675\ln(0.4200) \approx -0.8675
    • Latent Gaussian log-probability: log⁡p(u)=−12ln⁡(2πσ2)−(u−μ)22σ2=−12ln⁡(2π⋅0.25)−0=−12ln⁡(1.5708)≈−0.2258\log p(u) = -\frac{1}{2} \ln(2\pi \sigma^2) - \frac{(u - \mu)^2}{2\sigma^2} = -\frac{1}{2} \ln(2\pi \cdot 0.25) - 0 = -\frac{1}{2} \ln(1.5708) \approx -0.2258
    • Squashed corrected log-probability: log⁡π(a′∣s′)=log⁡p(u)−ln⁡(0.4200)=−0.2258−(−0.8675)=+0.6417\log \pi(a' | s') = \log p(u) - \ln(0.4200) = -0.2258 - (-0.8675) = +0.6417
  3. Compute Soft Bellman Target: Let current temperature α=0.2\alpha = 0.2, reward r=1.0r = 1.0, discount γ=0.99\gamma = 0.99. The twin target critics evaluate the pair (s′,a′)(s', a') as:

    Qθˉ1(s′,a′)=5.0,Qθˉ2(s′,a′)=4.2  ⟹  min⁡=4.2Q_{\bar{\theta}_1}(s', a') = 5.0, \quad Q_{\bar{\theta}_2}(s', a') = 4.2 \implies \min = 4.2

    Soft state value with entropy deduction:

    Vsoft(s′)=4.2−αlog⁡π(a′∣s′)=4.2−(0.2)(0.6417)=4.2−0.1283=4.0717V_{\text{soft}}(s') = 4.2 - \alpha \log \pi(a' | s') = 4.2 - (0.2)(0.6417) = 4.2 - 0.1283 = 4.0717

    Bellman target yy:

    y=r+γVsoft(s′)=1.0+0.99(4.0717)=1.0+4.0309=5.0309y = r + \gamma V_{\text{soft}}(s') = 1.0 + 0.99(4.0717) = 1.0 + 4.0309 = 5.0309
  4. Update Temperature α\alpha: Compare current negative log-prob with target entropy Hˉ=−1.0\bar{\mathcal{H}} = -1.0:

    log⁡π(a∣s)+Hˉ=0.6417+(−1.0)=−0.3583\log \pi(a | s) + \bar{\mathcal{H}} = 0.6417 + (-1.0) = -0.3583

    Entropy difference: −log⁡π(a∣s)−Hˉ=−0.6417−(−1.0)=+0.3583-\log \pi(a|s) - \bar{\mathcal{H}} = -0.6417 - (-1.0) = +0.3583. Because −log⁡π>Hˉ-\log \pi > \bar{\mathcal{H}}, the policy has more entropy than the target constraint demands. The gradient on log⁡α\log \alpha is positive, triggering a decrease in α\alpha to focus policy execution.

Code

import numpy as np

class SoftActorCriticStep:    """Implements core mathematics of Soft Actor-Critic (SAC) updates:
    Reparameterized squashed Gaussian sampling, Jacobian correction,    clipped double Q soft Bellman targets, and automatic temperature tuning.    """
    def __init__(        self,        action_dim: int = 1,        gamma: float = 0.99,        initial_alpha: float = 0.2,        eps_stabilizer: float = 1e-6,    ) -> None:        self.action_dim = action_dim        self.gamma = gamma        self.log_alpha = float(np.log(initial_alpha))        self.target_entropy = -float(action_dim)        self.eps = eps_stabilizer
    @property    def alpha(self) -> float:        return float(np.exp(self.log_alpha))
    def sample_action(        self,        mu: np.ndarray,        sigma: np.ndarray,        noise: np.ndarray | None = None,    ) -> tuple[np.ndarray, np.ndarray, float]:        """Samples action via reparameterization trick: u = mu + sigma * eps, a = tanh(u).
        Applies Jacobian change-of-variables correction to calculate log pi(a|s).        """        if noise is None:            noise = np.random.randn(*mu.shape)
        # Reparameterization trick: differentiable latent Gaussian sample        u = mu + sigma * noise        a = np.tanh(u)
        # Gaussian density in unconstrained latent space u        var = sigma**2        log_p_u = -0.5 * np.sum(np.log(2.0 * np.pi * var) + ((u - mu) ** 2) / var)
        # Jacobian change-of-variables correction for tanh squashing        log_jacobian = float(np.sum(np.log(1.0 - a**2 + self.eps)))        log_pi = float(log_p_u - log_jacobian)
        return u, a, log_pi
    def compute_soft_target(        self,        reward: float,        q1_target: float,        q2_target: float,        next_log_pi: float,    ) -> float:        """Computes soft Bellman value target with clipped double Q and entropy bonus."""        min_q = min(q1_target, q2_target)        soft_value = min_q - self.alpha * next_log_pi        y = reward + self.gamma * soft_value        return float(y)
    def compute_temperature_update(        self,        log_pi: float,        lr_alpha: float = 0.05,    ) -> tuple[float, float, float]:        """Computes automatic entropy temperature loss, gradient, and updated alpha."""        entropy_diff = -log_pi - self.target_entropy        alpha_loss = -self.log_alpha * (log_pi + self.target_entropy)        grad_log_alpha = entropy_diff
        # Gradient descent step on log(alpha)        new_log_alpha = self.log_alpha - lr_alpha * grad_log_alpha        new_alpha = float(np.exp(new_log_alpha))
        return alpha_loss, grad_log_alpha, new_alpha

if __name__ == "__main__":    np.set_printoptions(precision=4, suppress=True)
    # Initialize SAC step for 1D continuous action space    sac = SoftActorCriticStep(action_dim=1, gamma=0.99, initial_alpha=0.2)
    # Transition parameters matching the worked numerical example    mu = np.array([1.0], dtype=np.float64)    sigma = np.array([0.5], dtype=np.float64)    noise = np.array([0.0], dtype=np.float64)
    u, a, log_pi = sac.sample_action(mu, sigma, noise=noise)
    q1_next, q2_next = 5.0, 4.2    reward = 1.0    soft_target = sac.compute_soft_target(reward, q1_next, q2_next, log_pi)
    alpha_loss, grad_log_alpha, updated_alpha = sac.compute_temperature_update(        log_pi, lr_alpha=0.05    )
    print(f"Latent sample u: {u[0]:.4f}")    # -> Latent sample u: 1.0000
    print(f"Squashed action a: {a[0]:.4f}")    # -> Squashed action a: 0.7616
    print(f"Corrected log pi(a|s): {log_pi:.4f}")    # -> Corrected log pi(a|s): 0.6418
    print(f"Soft Bellman target y: {soft_target:.4f}")    # -> Soft Bellman target y: 5.0309
    print(f"Temperature alpha: {sac.alpha:.4f}")    # -> Temperature alpha: 0.2000
    print(f"Target entropy H_bar: {sac.target_entropy:.4f}")    # -> Target entropy H_bar: -1.0000
    print(f"Log-alpha gradient: {grad_log_alpha:.4f}")    # -> Log-alpha gradient: 0.3582
    print(f"Updated alpha: {updated_alpha:.4f}")    # -> Updated alpha: 0.1964
    # Assert correctness    assert np.isclose(a[0], 0.7616, atol=1e-3)    assert np.isclose(soft_target, 5.0309, atol=1e-3)    assert updated_alpha < sac.alpha  # Alpha reduced as entropy > target

Watch Out For

Omitting the Tanh Jacobian Change-of-Variables Correction

A common bug in custom SAC implementations is computing log⁡π(a∣s)\log \pi(a|s) using the raw Gaussian probability density function log⁡p(u∣s)\log p(u|s) without subtracting the Jacobian determinant penalty ∑ilog⁡(1−ai2+δ)\sum_i \log(1 - a_i^2 + \delta).

Because tanh⁡\tanh compresses (−∞,∞)(-\infty, \infty) into (−1,1)(-1, 1), probability mass piles up heavily near the action boundaries ±1\pm 1. Without the Jacobian correction, the policy's calculated entropy is vastly overestimated near the boundaries, leading the automatic temperature tuner to drive α→0\alpha \to 0 prematurely or causing severe gradient explosion during backpropagation. Always apply the change-of-variables formula with numerical stabilization δ\delta.

The Quick Version

  • Maximum Entropy Framework: SAC optimizes for both cumulative rewards and policy entropy, encouraging wide exploration and preventing convergence to brittle, deterministic policies.
  • Clipped Double Q-Learning: Uses the minimum prediction of two independent target critics (min⁡(Q1,Q2)\min(Q_1, Q_2)) in soft Bellman backups to suppress value overestimation.
  • Reparameterized Tanh Actor: Enables end-to-end backpropagation through stochastic action choices while bounding continuous actions to (−1,1)(-1, 1) with an analytical Jacobian correction.
  • Automatic Temperature Tuning: Continuously modulates the entropy weight α\alpha to satisfy an entropy constraint Hˉ=−dim⁡(A)\bar{\mathcal{H}} = -\dim(\mathcal{A}), balancing exploration early and exploitation late.