Skip to content
AI360Xpert
Beta

Circuit Discovery Methodologies

By surgically swapping intermediate activations between clean and corrupted prompts, circuit discovery isolates the exact sub-network of attention heads and neurons causally responsible for a model's prediction.

Activation patching workflow: swapping intermediate activations between clean and corrupted runs to isolate causal circuit edges.
Activation patching workflow: swapping intermediate activations between clean and corrupted runs to isolate causal circuit edges.

Why Does This Exist?

A 70-billion parameter Large Language Model executes trillions of floating-point operations for every single generated token. When the model correctly answers a grammar puzzle, tracks a variable across Python code, or executes Indirect Object Identification (e.g., "When Mary and John went to the store, John gave a drink to [Mary]"), looking at the full network is overwhelming. Millions of parameters are firing simultaneously.

However, empirical research by Wang et al. and the mechanistic interpretability community revealed that individual tasks do not require the whole model. Instead, specific reasoning behaviors are computed by a tiny, sparse sub-graph of nodes—often fewer than 30 attention heads and MLP layers out of thousands. The remaining 99.9% of the network acts as an identity pass-through for that task.

Traditional correlational methods (such as attention weights or probing classifiers) fail to find this sub-graph because correlated activations are not necessarily causal. A head might attend strongly to "Mary" simply because it echoes proper nouns everywhere, without transmitting that information to the final output logits. Circuit discovery methodologies provide rigorous causal interventional frameworks—such as activation patching, causal tracing, and Edge Attribution Patching (EAP)—to surgically isolate, verify, and document the minimal causal circuits responsible for specific model capabilities.

Think of It Like This

Swapping circuit board chips between working and broken radios

Imagine you possess two identical vintage radios: Radio A and Radio B. Radio A plays clean, crisp FM music. Radio B only outputs loud, staticky white noise.

You want to find which component on the circuit board is responsible for tuning the FM frequency. You do not probe every wire with an oscilloscope and guess from electrical hums.

Instead, you use an interventional chip-swapping protocol. While both radios are plugged into power, you take Chip #1 from Radio A and swap it into Radio B. If Radio B still produces static, Chip #1 is irrelevant. You repeat this across the board. Suddenly, when you transplant Chip #7 from Radio A into Radio B, Radio B's static vanishes and it instantly plays clear music.

By performing a controlled surgical transplant, you have proven with causal certainty that Chip #7 is the frequency demodulator, regardless of what the other 50 chips on the board are doing.

How It Actually Works

Activation Patching and Causal Mediation Metrics

Circuit discovery converts qualitative hypotheses into quantifiable causal experiments through activation patching (also known as interchange intervention or causal mediation analysis):

1. Define Contrastive Pairs

Create a clean prompt xcleanx_{\text{clean}} and a corrupted counterpart xcorruptx_{\text{corrupt}} that differs only by the target entity:

  • xcleanx_{\text{clean}}: "When Mary and John went to the store, John gave a drink to" →\to Target: Mary
  • xcorruptx_{\text{corrupt}}: "When Alice and Bob went to the store, Bob gave a drink to" →\to Target: Alice

We define the primary performance metric as the difference between the target logit and the distractor logit at the final token position:

Δlogit(x)=logit(Target)−logit(Distractor)\Delta \text{logit}(x) = \text{logit}(\text{Target}) - \text{logit}(\text{Distractor})

2. Run Forward Baselines

Run both prompts through the network without intervention and record internal activation caches:

  • Δlogitclean=logitclean(Mary)−logitclean(John)>0\Delta \text{logit}_{\text{clean}} = \text{logit}_{\text{clean}}(\text{Mary}) - \text{logit}_{\text{clean}}(\text{John}) > 0 (High positive score)
  • Δlogitcorrupt=logitcorrupt(Mary)−logitcorrupt(John)≤0\Delta \text{logit}_{\text{corrupt}} = \text{logit}_{\text{corrupt}}(\text{Mary}) - \text{logit}_{\text{corrupt}}(\text{John}) \le 0 (Negative or zero score)

3. Perform Activation Patching Interventions

To test whether node vv (an attention head or MLP layer at layer ll) is causally responsible for transmitting the correct entity:

  1. Pass xcorruptx_{\text{corrupt}} through the model.
  2. At layer ll, hook node vv's output and overwrite it with the cached activation from the clean run:
avpatched←avcleana_v^{\text{patched}} \leftarrow a_v^{\text{clean}}
  1. Allow all subsequent layers to execute normally using the patched activation.
  2. Measure the patched logit difference Δlogitpatched\Delta \text{logit}_{\text{patched}}.

4. Compute Normalized Causal Mediation Effect

We normalize the recovery score to measure what fraction of the clean behavior was restored solely by restoring node vv:

Normalized Effect(v)=Δlogitpatched−ΔlogitcorruptΔlogitclean−Δlogitcorrupt\text{Normalized Effect}(v) = \frac{\Delta \text{logit}_{\text{patched}} - \Delta \text{logit}_{\text{corrupt}}}{\Delta \text{logit}_{\text{clean}} - \Delta \text{logit}_{\text{corrupt}}}
  • Effect≈1.0\text{Effect} \approx 1.0: Node vv is a primary causal bottleneck carrying the essential information.
  • Effect≈0.0\text{Effect} \approx 0.0: Node vv is functionally inert for this task.

5. Scaled Circuit Discovery via Edge Attribution Patching (EAP)

Iterating activation patching over every individual edge in a network requires O(E)\mathcal{O}(E) forward passes, taking days of GPU compute. Modern discovery pipelines use Edge Attribution Patching (EAP), which linearizes the intervention using a single backward pass gradient:

Attribution(u→v)≈(auclean−aucorrupt)⋅∇auLcorrupt\text{Attribution}(u \to v) \approx \left( a_u^{\text{clean}} - a_u^{\text{corrupt}} \right) \cdot \nabla_{a_u} \mathcal{L}_{\text{corrupt}}

EAP evaluates all candidate edges across a 30-layer model in seconds, pruning the graph to leave only the verified minimal circuit.

Worked Example

Let us evaluate the causal importance of two candidate attention heads in a 12-layer language model executing the Indirect Object Identification (IOI) task.

  • Clean run logit difference: logit(Mary)=8.5\text{logit}(\text{Mary}) = 8.5, logit(John)=3.5  ⟹  Δlogitclean=8.5−3.5=+5.0\text{logit}(\text{John}) = 3.5 \implies \Delta \text{logit}_{\text{clean}} = 8.5 - 3.5 = +5.0
  • Corrupted run logit difference: logit(Mary)=2.0\text{logit}(\text{Mary}) = 2.0, logit(John)=4.0  ⟹  Δlogitcorrupt=2.0−4.0=−2.0\text{logit}(\text{John}) = 4.0 \implies \Delta \text{logit}_{\text{corrupt}} = 2.0 - 4.0 = -2.0
  • Denominator (total task swing): Δlogitclean−Δlogitcorrupt=5.0−(−2.0)=7.0\Delta \text{logit}_{\text{clean}} - \Delta \text{logit}_{\text{corrupt}} = 5.0 - (-2.0) = 7.0

Test Candidate A: Attention Head Layer 9, Head 6 (L9H6) We execute the corrupted prompt, but freeze and patch the output of L9H6 with its clean activation:

  • Patched output logits: logit(Mary)=6.2\text{logit}(\text{Mary}) = 6.2, logit(John)=2.0  ⟹  Δlogitpatched=6.2−2.0=+4.2\text{logit}(\text{John}) = 2.0 \implies \Delta \text{logit}_{\text{patched}} = 6.2 - 2.0 = +4.2
  • Compute Normalized Causal Effect:
Normalized Effect(L9H6)=4.2−(−2.0)7.0=6.27.0≈0.8857(88.6%)\text{Normalized Effect}(\text{L9H6}) = \frac{4.2 - (-2.0)}{7.0} = \frac{6.2}{7.0} \approx 0.8857 \quad (88.6\%)

Transplanting this single attention head alone restores 88.6% of the model's clean capability. L9H6 is verified as a primary Name Mover Head.

Test Candidate B: Attention Head Layer 4, Head 1 (L4H1) We repeat the experiment, patching L4H1 instead:

  • Patched output logits: logit(Mary)=2.1\text{logit}(\text{Mary}) = 2.1, logit(John)=3.9  ⟹  Δlogitpatched=2.1−3.9=−1.8\text{logit}(\text{John}) = 3.9 \implies \Delta \text{logit}_{\text{patched}} = 2.1 - 3.9 = -1.8
  • Compute Normalized Causal Effect:
Normalized Effect(L4H1)=−1.8−(−2.0)7.0=0.27.0≈0.0286(2.9%)\text{Normalized Effect}(\text{L4H1}) = \frac{-1.8 - (-2.0)}{7.0} = \frac{0.2}{7.0} \approx 0.0286 \quad (2.9\%)

Transplanting L4H1 only recovers 2.9% of the metric. L4H1 is causally pruned from the circuit.

Code

Below is a self-contained Python implementation demonstrating an activation patching harness on a toy computation graph:

import numpy as np

class ToyTransformerNode:    """Simulates a modular attention head or MLP layer."""
    def __init__(self, weight: float, bias: float) -> None:        self.weight = weight        self.bias = bias
    def forward(self, x: float) -> float:        return self.weight * x + self.bias

class ToyGraphModel:    """Toy 3-stage model: Input -> [Node A, Node B] -> Final Logit Head."""
    def __init__(self) -> None:        self.node_a = ToyTransformerNode(weight=1.8, bias=0.2)  # Causal carrier        self.node_b = ToyTransformerNode(weight=0.1, bias=0.0)  # Distractor node
    def run(self, x: float, patch_a: float | None = None, patch_b: float | None = None) -> tuple[float, float, float]:        act_a = self.node_a.forward(x)        act_b = self.node_b.forward(x)
        # Apply surgical interventional patches if present        if patch_a is not None:            act_a = patch_a        if patch_b is not None:            act_b = patch_b
        # Final readout layer combining activations        output_logit = 2.0 * act_a + 0.5 * act_b        return output_logit, act_a, act_b

model = ToyGraphModel()
# Step 1: Clean and Corrupted Baseline Runsx_clean, x_corrupt = 2.0, 0.5
logit_clean, clean_a, clean_b = model.run(x_clean)logit_corrupt, corrupt_a, corrupt_b = model.run(x_corrupt)total_swing = logit_clean - logit_corrupt
# Step 2: Patch Node A into corrupted runpatched_a_logit, _, _ = model.run(x_corrupt, patch_a=clean_a)effect_a = (patched_a_logit - logit_corrupt) / total_swing
# Step 3: Patch Node B into corrupted runpatched_b_logit, _, _ = model.run(x_corrupt, patch_b=clean_b)effect_b = (patched_b_logit - logit_corrupt) / total_swing
print(f"Clean Logit: {logit_clean:.2f} | Corrupt Logit: {logit_corrupt:.2f}")# -> Clean Logit: 7.65 | Corrupt Logit: 2.22print(f"Node A Causal Recovery: {effect_a * 100:.1f}%")# -> Node A Causal Recovery: 98.4%print(f"Node B Causal Recovery: {effect_b * 100:.1f}%")# -> Node B Causal Recovery: 1.6%

Watch Out For

Self-repair and backup 'Hydra heads' masking causal circuits during ablation

When isolating circuits, practitioners often perform zero-ablation (knocking out a head by setting its output tensor to zero). Frequently, they observe that zeroing out a critical attention head barely degrades the final output logit.

This is the self-repair trap. Large models develop backup redundancies called "Hydra heads". When a primary attention head is zero-ablated, downstream heads in later layers detect the missing activation in the residual stream and immediately increase their own output magnitudes to compensate, masking the causal importance of the knocked-out head.

Fix: Never rely solely on zero-ablation to determine circuit boundaries. Use resample patching or denoising activation patching (swapping activations from an in-distribution corrupted run rather than setting tensors to zero). Corrupted activations maintain natural token distributions and avoid triggering the model's anomalous self-repair circuitry.

The Quick Version

  • Circuit discovery identifies the minimal subgraphs of attention heads and MLP layers causally responsible for specific model capabilities.
  • Activation patching replaces internal activations during a corrupted run with counterparts from a clean run to measure performance recovery.
  • Normalized causal mediation metrics quantify the exact percentage of behavior attributable to individual nodes or edges.
  • Edge Attribution Patching (EAP) scales discovery across deep models by linearizing interventions using single-pass gradient attributions.