Skip to content
AI360Xpert
Beta

Noisy Nets for Exploration

Noisy Nets inject learnable parametric noise into neural network weights, driving state-dependent exploration that automatically self-anneals as learning progresses.

Noisy Nets decompose network weights into learnable mean and scale parameters, replacing action dithering with structured, self-annealing exploration in parameter space.
Noisy Nets decompose network weights into learnable mean and scale parameters, replacing action dithering with structured, self-annealing exploration in parameter space.

Why Does This Exist?

In classic deep reinforcement learning, agents rely on heuristic ϵ\epsilon-greedy exploration: with probability 1−ϵ1 - \epsilon, the agent selects the greedy action arg⁡max⁡aQ(s,a)\arg\max_a Q(s, a), and with probability ϵ\epsilon, it samples a uniformly random action from A\mathcal{A}.

While simple, ϵ\epsilon-greedy suffers from three critical deficiencies:

  1. Destructive Action Dithering: ϵ\epsilon-greedy explores in action space by sampling independent random choices at every single time step. This introduces Brownian random-walk dynamics: the agent moves left, then right, then left, canceling its own progress. In environments requiring coherent multi-step navigation (such as navigating a maze or solving sparse-reward games like Montezuma's Revenge), random dithering virtually never reaches distant goals.
  2. State-Agnostic Exploration: ϵ\epsilon-greedy treats every state identically. Even after an agent has completely mastered opening moves, it continues executing random, game-ending blunders with probability ϵ\epsilon.
  3. Brittle Hyperparameter Schedules: Designing an exploration schedule requires manually tuning ϵstart\epsilon_{\text{start}}, ϵend\epsilon_{\text{end}}, and decay rates across millions of frames, which rarely transfers across different tasks.

Introduced by Meire Fortunato et al. (DeepMind, 2018), Noisy Networks for Exploration (Noisy Nets) replace heuristic action dithering with learnable parametric noise injected directly into the neural network's weights. By perturbing parameters rather than actions, the policy remains temporally consistent across consecutive decisions, enabling structured, deep exploration.

Crucially, because the noise scale parameters σ\sigma are updated by gradient descent, exploration naturally self-anneals: as value uncertainty declines in mastered states, gradients automatically drive σ→0\sigma \to 0.

Think of It Like This

The Trembling Hands of a Novice Musician

Imagine an aspiring student learning to play acoustic guitar or violin.

If the student practiced using ϵ\epsilon-greedy exploration, they would play nine notes with textbook precision, but on every tenth note, an involuntary muscle spasm would cause them to strike an unrelated string or drop the bow entirely. That abrupt action dithering does not teach acoustic subtlety—it merely derails the melody and destroys musical rhythm.

With Noisy Nets (parametric weight noise), the exploration originates internally within the motor cortex. During early practice, the student's grip, wrist tension, and finger pressure are slightly unsteady (σ\sigma is large). This continuous tremor subtly alters vibrato, fretboard angle, and string contact across entire musical phrases. Because the variation is integrated directly into their motor coordination, the student discovers resonant harmonics, alternative fingerings, and dynamic expression.

As muscle memory consolidates and auditory feedback confirms pitch accuracy, the motor cortex naturally tightens control: the tremor variance σ\sigma dampens toward zero on mastered chords. However, when the student encounters an unfamiliar, complex chord progression in a minor key, flexibility naturally re-emerges to search for optimal finger placement.

Where the analogy stops: Human muscle tremor is physical biomechanics. In Noisy Nets, the tremor is precisely factorised Gaussian noise sampled on network forward passes and modulated by mathematical backpropagation gradients.

How It Actually Works

Mathematical Formulation and Factorised Noise Architecture

In a standard deterministic fully-connected layer, an input vector x∈Rpx \in \mathbb{R}^p is transformed to output y∈Rqy \in \mathbb{R}^q via:

y=Wx+by = W x + b

Where W∈Rq×pW \in \mathbb{R}^{q \times p} and b∈Rqb \in \mathbb{R}^q are learnable parameter tensors.

A NoisyLinear layer replaces fixed weights and biases with random variables parameterized by learnable means (μ\mu) and learnable noise scale deviations (σ\sigma):

y≐(μw+σw⊙εw)x+(μb+σb⊙εb)y \doteq \left( \mu^w + \sigma^w \odot \varepsilon^w \right) x + \left( \mu^b + \sigma^b \odot \varepsilon^b \right)

Where:

  • μw∈Rq×p\mu^w \in \mathbb{R}^{q \times p} and μb∈Rq\mu^b \in \mathbb{R}^q are learnable mean parameters.
  • σw∈Rq×p\sigma^w \in \mathbb{R}^{q \times p} and σb∈Rq\sigma^b \in \mathbb{R}^q are learnable noise scale parameters.
  • εw∈Rq×p\varepsilon^w \in \mathbb{R}^{q \times p} and εb∈Rq\varepsilon^b \in \mathbb{R}^q are zero-mean stochastic noise tensors.
  • ⊙\odot denotes element-wise (Hadamard) multiplication.

Independent vs. Factorised Gaussian Noise

There are two primary ways to generate the noise tensors εw\varepsilon^w and εb\varepsilon^b:

  1. Independent Gaussian Noise: Every element εi,jw\varepsilon^w_{i, j} and εib\varepsilon^b_i is sampled independently from N(0,1)\mathcal{N}(0, 1). This requires generating p×q+qp \times q + q random variables per forward pass, which becomes computationally prohibitive in deep networks with wide hidden layers.

  2. Factorised Gaussian Noise (Standard in Deep RL): Instead of generating p×qp \times q independent variables, factorised noise generates only pp unit Gaussian variables for the inputs (εin∈Rp\varepsilon_{\text{in}} \in \mathbb{R}^p) and qq variables for the outputs (εout∈Rq\varepsilon_{\text{out}} \in \mathbb{R}^q), requiring only p+qp + q random draws.

Each component is transformed using the non-linear real-valued function:

f(z)≐sgn⁡(z)∣z∣f(z) \doteq \operatorname{sgn}(z) \sqrt{|z|}

The weight and bias noise matrices are then constructed via outer product:

εi,jw≐f(εout,i)⋅f(εin,j),εib≐f(εout,i)\varepsilon^w_{i, j} \doteq f(\varepsilon_{\text{out}, i}) \cdot f(\varepsilon_{\text{in}, j}), \quad \varepsilon^b_i \doteq f(\varepsilon_{\text{out}, i})

The Self-Annealing Gradient Mechanism

During backpropagation, gradients flow into both the mean parameters and the scale parameters:

∂L∂μi,jw=∂L∂yixj,∂L∂σi,jw=∂L∂yixj⋅εi,jw\frac{\partial \mathcal{L}}{\partial \mu^w_{i, j}} = \frac{\partial \mathcal{L}}{\partial y_i} x_j, \quad \frac{\partial \mathcal{L}}{\partial \sigma^w_{i, j}} = \frac{\partial \mathcal{L}}{\partial y_i} x_j \cdot \varepsilon^w_{i, j}

∂L∂μib=∂L∂yi,∂L∂σib=∂L∂yi⋅εib\frac{\partial \mathcal{L}}{\partial \mu^b_i} = \frac{\partial \mathcal{L}}{\partial y_i}, \quad \frac{\partial \mathcal{L}}{\partial \sigma^b_i} = \frac{\partial \mathcal{L}}{\partial y_i} \cdot \varepsilon^b_i

Why does exploration self-anneal? Consider a state where the agent has already converged to the correct value function. In this regime, weight perturbations σw⊙εw\sigma^w \odot \varepsilon^w cause the network's output to deviate from the optimal Bellman target, increasing the loss L\mathcal{L}. Because the gradient ∂L∂σi,jw\frac{\partial \mathcal{L}}{\partial \sigma^w_{i, j}} is positively correlated with loss increases, gradient descent pushes σi,jw\sigma^w_{i, j} downward toward zero.

In contrast, in novel states with high value uncertainty, exploratory perturbations uncover rewarding trajectories that reduce the loss, sustaining non-zero σ\sigma values. The network autonomously decides where and how much to explore without manual schedules.


Worked Numerical Example

Consider a single NoisyLinear neuron with p=2p = 2 inputs and q=1q = 1 output.

Step 1: Initial Parameters and Input

  • Input vector: x=[2.0,−1.0]⊤x = [2.0, -1.0]^\top
  • Mean parameters: μw=[0.5,−0.2]\mu^w = [0.5, -0.2], μb=0.1\mu^b = 0.1
  • Scale parameters: σw=[0.4,0.3]\sigma^w = [0.4, 0.3], σb=0.2\sigma^b = 0.2
  • Learning rate: α=0.1\alpha = 0.1

Step 2: Factorised Noise Generation

  • Sample input noise: εin=[1.0,−4.0]\varepsilon_{\text{in}} = [1.0, -4.0]
  • Sample output noise: εout=[1.0]\varepsilon_{\text{out}} = [1.0]

Apply transform f(z)=sgn⁡(z)∣z∣f(z) = \operatorname{sgn}(z)\sqrt{|z|}:

  • f(εin[0])=f(1.0)=sgn⁡(1.0)1.0=1.0f(\varepsilon_{\text{in}}[0]) = f(1.0) = \operatorname{sgn}(1.0)\sqrt{1.0} = 1.0
  • f(εin[1])=f(−4.0)=sgn⁡(−4.0)4.0=−2.0f(\varepsilon_{\text{in}}[1]) = f(-4.0) = \operatorname{sgn}(-4.0)\sqrt{4.0} = -2.0
  • f(εout[0])=f(1.0)=sgn⁡(1.0)1.0=1.0f(\varepsilon_{\text{out}}[0]) = f(1.0) = \operatorname{sgn}(1.0)\sqrt{1.0} = 1.0

Compute noise matrices:

  • ε0w=f(εout[0])⋅f(εin[0])=1.0×1.0=1.0\varepsilon^w_0 = f(\varepsilon_{\text{out}}[0]) \cdot f(\varepsilon_{\text{in}}[0]) = 1.0 \times 1.0 = 1.0
  • ε1w=f(εout[0])⋅f(εin[1])=1.0×(−2.0)=−2.0\varepsilon^w_1 = f(\varepsilon_{\text{out}}[0]) \cdot f(\varepsilon_{\text{in}}[1]) = 1.0 \times (-2.0) = -2.0
  • εb=f(εout[0])=1.0\varepsilon^b = f(\varepsilon_{\text{out}}[0]) = 1.0

Step 3: Realized Weights and Forward Pass

Realize perturbed parameters: w0=μ0w+σ0wε0w=0.5+0.4×(1.0)=0.90w_0 = \mu^w_0 + \sigma^w_0 \varepsilon^w_0 = 0.5 + 0.4 \times (1.0) = 0.90 w1=μ1w+σ1wε1w=−0.2+0.3×(−2.0)=−0.2−0.6=−0.80w_1 = \mu^w_1 + \sigma^w_1 \varepsilon^w_1 = -0.2 + 0.3 \times (-2.0) = -0.2 - 0.6 = -0.80 b=μb+σbεb=0.1+0.2×(1.0)=0.30b = \mu^b + \sigma^b \varepsilon^b = 0.1 + 0.2 \times (1.0) = 0.30

Compute perturbed forward output: y=w0x0+w1x1+b=(0.90×2.0)+(−0.80×−1.0)+0.30=1.80+0.80+0.30=2.90y = w_0 x_0 + w_1 x_1 + b = (0.90 \times 2.0) + (-0.80 \times -1.0) + 0.30 = 1.80 + 0.80 + 0.30 = 2.90

(Note: The deterministic unperturbed mean output is μw⋅x+μb=1.0+0.2+0.1=1.30\mu^w \cdot x + \mu^b = 1.0 + 0.2 + 0.1 = 1.30. Parametric noise shifted the output by +1.60+1.60.)

Step 4: Loss and Gradient Backpropagation

Suppose the target Bellman value is y∗=1.50y^* = 1.50.

  • Loss: L=12(y−y∗)2=12(2.90−1.50)2=0.98\mathcal{L} = \frac{1}{2} (y - y^*)^2 = \frac{1}{2} (2.90 - 1.50)^2 = 0.98
  • Output error: δ=y−y∗=2.90−1.50=1.40\delta = y - y^* = 2.90 - 1.50 = 1.40

Compute parameter gradients:

  • ∇μwL=δ⋅x=1.40×[2.0,−1.0]=[2.80,−1.40]\nabla_{\mu^w} \mathcal{L} = \delta \cdot x = 1.40 \times [2.0, -1.0] = [2.80, -1.40]
  • ∇σwL=δ⋅(x⊙εw)=1.40×[2.0×1.0,−1.0×(−2.0)]=[2.80,2.80]\nabla_{\sigma^w} \mathcal{L} = \delta \cdot (x \odot \varepsilon^w) = 1.40 \times [2.0 \times 1.0, -1.0 \times (-2.0)] = [2.80, 2.80]
  • ∇μbL=δ=1.40\nabla_{\mu^b} \mathcal{L} = \delta = 1.40
  • ∇σbL=δ⋅εb=1.40×1.0=1.40\nabla_{\sigma^b} \mathcal{L} = \delta \cdot \varepsilon^b = 1.40 \times 1.0 = 1.40

Step 5: SGD Parameter Update (α=0.1\alpha = 0.1)

μw←[0.5,−0.2]−0.1×[2.80,−1.40]=[0.22,−0.06]\mu^w \leftarrow [0.5, -0.2] - 0.1 \times [2.80, -1.40] = [0.22, -0.06] σw←[0.4,0.3]−0.1×[2.80,2.80]=[0.12,0.02]\sigma^w \leftarrow [0.4, 0.3] - 0.1 \times [2.80, 2.80] = [0.12, 0.02] μb←0.1−0.1×1.40=−0.04\mu^b \leftarrow 0.1 - 0.1 \times 1.40 = -0.04 σb←0.2−0.1×1.40=0.06\sigma^b \leftarrow 0.2 - 0.1 \times 1.40 = 0.06

The Self-Annealing Effect in Action: Because the positive noise perturbation caused the network to significantly overshoot the target (y=2.90y = 2.90 vs. y∗=1.50y^* = 1.50), gradient descent directly reduced the noise scale parameters:

  • σ0w\sigma^w_0 decreased from 0.40→0.120.40 \to 0.12
  • σ1w\sigma^w_1 decreased from 0.30→0.020.30 \to 0.02
  • σb\sigma^b decreased from 0.20→0.060.20 \to 0.06

The layer autonomously dampened its noise parameters to restore precision.

Code

import mathimport randomfrom typing import List, Sequence, Tuple

def f_transform(x: float) -> float:    """Factorised noise non-linear transform: f(x) = sgn(x) * sqrt(|x|)."""    return math.copysign(math.sqrt(abs(x)), x)

class NoisyLinear:    """Fully-connected layer with factorised Gaussian parametric noise."""
    def __init__(self, in_features: int, out_features: int, sigma_zero: float = 0.5) -> None:        self.in_features = in_features        self.out_features = out_features
        # Initialize mean parameters (mu) and scale parameters (sigma)        bound = 1.0 / math.sqrt(in_features)        self.mu_w: List[List[float]] = [            [random.uniform(-bound, bound) for _ in range(in_features)]            for _ in range(out_features)        ]        self.mu_b: List[float] = [random.uniform(-bound, bound) for _ in range(out_features)]
        sigma_w_init = sigma_zero / math.sqrt(in_features)        self.sigma_w: List[List[float]] = [            [sigma_w_init for _ in range(in_features)] for _ in range(out_features)        ]        self.sigma_b: List[float] = [            sigma_zero / math.sqrt(out_features) for _ in range(out_features)        ]
        # Internal noise buffers        self.eps_in: List[float] = [0.0] * in_features        self.eps_out: List[float] = [0.0] * out_features
    def sample_noise(        self,        eps_in: Sequence[float] | None = None,        eps_out: Sequence[float] | None = None,    ) -> None:        """Sample or inject factorised Gaussian noise."""        if eps_in is not None and eps_out is not None:            self.eps_in = list(eps_in)            self.eps_out = list(eps_out)        else:            self.eps_in = [random.gauss(0.0, 1.0) for _ in range(self.in_features)]            self.eps_out = [random.gauss(0.0, 1.0) for _ in range(self.out_features)]
    def forward(self, x: Sequence[float]) -> List[float]:        """Compute y = (mu_w + sigma_w * eps_w) x + (mu_b + sigma_b * eps_b)."""        f_in = [f_transform(v) for v in self.eps_in]        f_out = [f_transform(v) for v in self.eps_out]
        outputs = [0.0] * self.out_features        for i in range(self.out_features):            b_realized = self.mu_b[i] + self.sigma_b[i] * f_out[i]            w_dot_x = 0.0            for j in range(self.in_features):                weight_noise = f_out[i] * f_in[j]                w_realized = self.mu_w[i][j] + self.sigma_w[i][j] * weight_noise                w_dot_x += w_realized * x[j]            outputs[i] = w_dot_x + b_realized
        return outputs

if __name__ == "__main__":    # Reproduce the exact worked numerical example (p=2 inputs, q=1 output)    layer = NoisyLinear(in_features=2, out_features=1)
    # Set parameters to exact test values    layer.mu_w[0] = [0.5, -0.2]    layer.mu_b[0] = 0.1    layer.sigma_w[0] = [0.4, 0.3]    layer.sigma_b[0] = 0.2
    # Inject exact factorised noise vectors    layer.sample_noise(eps_in=[1.0, -4.0], eps_out=[1.0])
    x = [2.0, -1.0]    y = layer.forward(x)[0]    print(f"Forward Output y: {y:.4f}")
    # Compute loss against target y* = 1.50    y_star = 1.50    delta = y - y_star    loss = 0.5 * (delta**2)    print(f"Target: {y_star:.2f}, Delta: {delta:.4f}, Loss: {loss:.4f}")
    # Compute analytical gradients    f_in = [f_transform(v) for v in layer.eps_in]    f_out = [f_transform(v) for v in layer.eps_out]    eps_w = [f_out[0] * f_in[0], f_out[0] * f_in[1]]    eps_b = f_out[0]
    grad_mu_w = [delta * x[j] for j in range(2)]    grad_sigma_w = [delta * x[j] * eps_w[j] for j in range(2)]    grad_mu_b = delta    grad_sigma_b = delta * eps_b
    print(f"Grad mu_w:    {[round(g, 4) for g in grad_mu_w]}")    print(f"Grad sigma_w: {[round(g, 4) for g in grad_sigma_w]}")    print(f"Grad mu_b:    {grad_mu_b:.4f}")    print(f"Grad sigma_b: {grad_sigma_b:.4f}")
    # Apply SGD parameter update with learning rate alpha = 0.1    lr = 0.1    layer.mu_w[0] = [layer.mu_w[0][j] - lr * grad_mu_w[j] for j in range(2)]    layer.sigma_w[0] = [layer.sigma_w[0][j] - lr * grad_sigma_w[j] for j in range(2)]    layer.mu_b[0] -= lr * grad_mu_b    layer.sigma_b[0] -= lr * grad_sigma_b
    print("\nUpdated Parameters after 1 Step:")    print(f"mu_w:    {[round(v, 4) for v in layer.mu_w[0]]}")    print(f"sigma_w: {[round(v, 4) for v in layer.sigma_w[0]]}")    print(f"mu_b:    {layer.mu_b[0]:.4f}")    print(f"sigma_b: {layer.sigma_b[0]:.4f}")
# Expected Output:# Forward Output y: 2.9000# Target: 1.50, Delta: 1.4000, Loss: 0.9800# Grad mu_w:    [2.8, -1.4]# Grad sigma_w: [2.8, 2.8]# Grad mu_b:    1.4000# Grad sigma_b: 1.4000# # Updated Parameters after 1 Step:# mu_w:    [0.22, -0.06]# sigma_w: [0.12, 0.02]# mu_b:    -0.0400# sigma_b: 0.0600

Watch Out For

The Step-Level Noise Resampling Trap

A frequent mistake when integrating Noisy Nets into deep Q-networks is resampling the noise tensors ε\varepsilon at every single environment decision step.

If you draw fresh noise vectors on every individual time step, parameter space noise collapses into high-frequency Brownian dithering. The agent acts on a different set of perturbed weights every fraction of a second, destroying the temporally extended trajectories that make Noisy Nets effective.

The Fix: Resample the noise tensors εin\varepsilon_{\text{in}} and εout\varepsilon_{\text{out}} strictly once per environment step (or once per episode in actor-critic setups), keeping the realized weights frozen while computing Q-values across all discrete action candidates a∈Aa \in \mathcal{A} for the current state. During training, sample a new set of noise for each sampled mini-batch to compute the Bellman error gradient.

The Quick Version

  • Weight Space Exploration: Replaces heuristic ϵ\epsilon-greedy action dithering by injecting stochastic noise directly into neural network weights (y=(μw+σw⊙εw)x+(μb+σb⊙εb)y = (\mu^w + \sigma^w \odot \varepsilon^w)x + (\mu^b + \sigma^b \odot \varepsilon^b)).
  • Temporally Consistent Policies: Because parameters remain stable across decision evaluations, the agent executes coherent multi-step exploratory sequences rather than random Brownian jitter.
  • Autonomous Self-Annealing: Gradients naturally penalize noise scale parameters σ\sigma in mastered states where perturbations increase Bellman loss, driving σ→0\sigma \to 0 without manual decay schedules.
  • Factorised Efficiency: Factorised Gaussian noise generates weight perturbations via outer products of input and output vectors (f(εout)f(εin)⊤f(\varepsilon_{\text{out}}) f(\varepsilon_{\text{in}})^\top), reducing noise variables from O(p⋅q)O(p \cdot q) down to O(p+q)O(p + q).