Skip to content
AI360Xpert
Beta

Feudal Networks (FuN)

A high-level Manager sets directional subgoals in a learned latent space every few steps, while a fast Worker learns primitive actions driven purely by cosine similarity to those goals.

FeUdal Networks decouple control into a dilated Manager setting directional latent subgoals and a fast Worker executing primitive actions driven by cosine-similarity intrinsic rewards.
FeUdal Networks decouple control into a dilated Manager setting directional latent subgoals and a fast Worker executing primitive actions driven by cosine-similarity intrinsic rewards.

Why Does This Exist?

In long-horizon reinforcement learning tasks with sparse rewards (such as the infamous Atari game Montezuma's Revenge), standard "flat" RL algorithms fail completely. When an agent must execute thousands of primitive actions before receiving a single scalar reward, credit assignment breaks down: the probability of discovering a rewarding trajectory through random exploratory motor noise is exponentially vanishing O(∣A∣−T)\mathcal{O}(|\mathcal{A}|^{-T}).

Hierarchical Reinforcement Learning (HRL) aims to solve this by decomposing complex behaviors into multi-level hierarchies. However, early HRL frameworks faced crippling bottlenecks:

  • Handcrafted Subgoals: Traditional architectures required engineers to manually define subgoals (e.g., discrete room coordinates or specific inventory flags), destroying end-to-end learning.
  • The Options Dilemma: Methods like the Options Framework required learning when to initiate and terminate options, suffering from option collapse where options degraded into single-step primitive actions.
  • Information Hiding Bottlenecks: Classical Feudal RL (Dayan & Hinton, 1993) hid state details from higher levels, creating severe communication bandwidth limits.

FeUdal Networks (FuN), introduced by Alexander Vezhnevets et al. at DeepMind in 2017, modernized feudal reinforcement learning for deep neural networks. FuN separates control into two distinct recurrent tiers operating at different temporal resolutions: a Manager operating at a dilated macro-timescale cc, and a Worker operating at primitive timescale t=1t=1.

Crucially, FuN eliminated manual subgoal engineering by defining subgoals as directional unit vectors in a learned continuous latent space. The Worker is motivated purely by cosine-similarity intrinsic rewards, while the Manager is trained using a novel Transition Policy Gradient that eliminates the need to backpropagate gradients through the Worker's motor decisions.

Think of It Like This

The Medieval Feudal Hierarchy

Imagine the administrative structure of a medieval agrarian kingdom:

The Feudal Lord (Manager) sits in the high castle tower observing kingdom-wide crop yields, seasonal rainfall, and long-range logistics. The Lord summons the local Baron (Worker) and issues an executive directional decree:

"Expand our agricultural territory toward the eastern valley."

The Lord does not know how to dig irrigation ditches, forge iron hoes, or swing scythes, and does not micromanage which serf plows which furrow. The Lord simply specifies the high-level directional objective gtg_t in strategic coordinates.

The Baron (Worker) lives on the ground with the farmers. The Baron receives personal wealth and prestige (Intrinsic Reward rtIr^I_t) proportionally to how effectively the workforce clears land in that commanded eastward direction (cos⁡(Δz,gt)\cos(\Delta z, g_t)). The Baron does not need to worry about kingdom-wide trade treaties or royal treasury reserves (Extrinsic Reward RtR_t)—they focus purely on solving local physical obstacles (boulders, rivers, mud) to satisfy the Lord's directional vector.

At harvest time, if the kingdom flourishes (RtR_t is high), the Lord receives royal credit. The Lord updates their strategic playbook (Transition Policy Gradient) to issue similar successful territorial commands in future seasons.

Where the analogy stops: In human feudalism, lords and barons communicate through spoken language that can be misunderstood, and information is often hidden. In FuN, both Manager and Worker share access to the same learned perceptual state ztz_t (the Non-Hiding Principle), ensuring hierarchical abstraction without communication bottlenecks.

How It Actually Works

Architectural Decoupling: Manager and Worker

FuN takes raw environment observations sts_t (such as visual image frames) and passes them through a shared convolutional perception backbone to produce a low-dimensional latent state:

zt=fpercept(st)∈Rdz_t = f_{\text{percept}}(s_t) \in \mathbb{R}^d

The architecture then splits into two specialized recurrent modules:

Observation s_t ──> [ Perception Backbone ] ──> Latent Representation z_t                                                     │         ┌───────────────────────────────────────────┴───────────────────────────────────────────┐         ▼                                                                                       ▼┌─────────────────────────────────────────┐                                     ┌─────────────────────────────────────────┐│ Dilated Manager RNN (Macro-step c)      │                                     │ Fast Worker RNN (Micro-step t = 1)      ││ Ticks every c steps (e.g., c = 10)      │                                     │ Ticks every single environment step     ││ Outputs Directional Goal g_t ∈ ℝ^d      │                                     │ Takes z_t and Goal Embedding            │└─────────────────────────────────────────┘                                     └─────────────────────────────────────────┘         │                                                                                       │         ▼                                                                                       ▼   Directional Goal g_t                                                            Primitive Motor Action a_t         │                                                                                       │         └───────────────────────────┬───────────────────────────────────────────────────────────┘                                     ▼                ┌─────────────────────────────────────────┐                │ Latent Cosine Alignment Module          │                │ Measures: cos(z_t - z_{t-i}, g_{t-i})   │                │ Yields: Worker Intrinsic Reward r^I_t   │                └─────────────────────────────────────────┘

1. The Dilated Manager (Macro-timescale cc)

The Manager operates at a coarse temporal resolution, ticking once every cc steps (typically c=10c = 10). To maintain temporal memory across hundreds of steps without vanishing gradients, the Manager uses a Dilated LSTM.

At macro-step tt, the Manager outputs a continuous latent goal vector gt∈Rdg_t \in \mathbb{R}^d, normalized to a unit vector:

g^t=gt∥gt∥\hat{g}_t = \frac{g_t}{\|g_t\|}

Key Insight: gtg_t is not an absolute target coordinate in state space. It is a directional displacement vector, commanding: "Over the next cc steps, shift the environment's latent state in direction g^t\hat{g}_t."

2. The Fast Worker (Primitive timescale t=1t=1)

The Worker operates at every single environment step tt. It receives the perceptual state ztz_t and an internal goal embedding wt=ϕ(gt)w_t = \phi(g_t). It produces a probability distribution over primitive motor actions:

at∼πworker(at∣zt,wt)a_t \sim \pi_{\text{worker}}\left(a_t \mid z_t, w_t\right)

The Worker handles low-level agility, collision avoidance, and motor reflexes.

Worker Intrinsic Reward via Cosine Similarity

The Worker is completely shielded from external environment rewards RtR_t. Instead, it is motivated purely by an internal cosine-similarity intrinsic reward rtIr^I_t.

At each step tt, the intrinsic reward measures how closely the actual state displacement achieved over the past cc steps aligns with the Manager's commanded directional goal:

rtI=1c∑i=1ccos⁡(zt−zt−i, gt−i)r^I_t = \frac{1}{c} \sum_{i=1}^c \cos\left(z_t - z_{t-i},\, g_{t-i}\right)

where the cosine similarity between two vectors uu and vv is:

cos⁡(u,v)=u⋅v∥u∥∥v∥\cos(u, v) = \frac{u \cdot v}{\|u\| \|v\|}

If the Worker's primitive actions move the agent in the exact direction commanded by the Manager, cos⁡(Δz,g)=+1.0\cos(\Delta z, g) = +1.0, maximizing intrinsic payoff. If the Worker moves perpendicular to the goal, reward is 0.00.0; if it moves in the opposite direction, reward is −1.0-1.0.

The Worker maximizes rtIr^I_t using standard policy gradient algorithms (such as A3C or PPO).

Transition Policy Gradients for the Manager

Training a hierarchical manager usually creates a mathematical dilemma: how can the Manager receive gradients without backpropagating through hundreds of unrolled Worker actions?

Vezhnevets et al. introduced Transition Policy Gradients by exploiting the structure of the goal space. The Worker is actively trained to align state displacements zt+c−ztz_{t+c} - z_t with goal gtg_t. Therefore, from the Manager's perspective, the probability of the environment transitioning from ztz_t to zt+cz_{t+c} can be modeled as:

p(zt+c∣zt,gt)∝exp⁡(cos⁡(zt+c−zt, gt))p\left(z_{t+c} \mid z_t, g_t\right) \propto \exp\left( \cos\left(z_{t+c} - z_t,\, g_t\right) \right)

Applying the policy gradient theorem to this transition distribution yields the Manager's update rule:

∇gtJM=E[(Rt−VM(st))∇gtcos⁡(zt+c−zt, gt)]\nabla_{g_t} J_M = \mathbb{E} \left[ \left(R_t - V_M(s_t)\right) \nabla_{g_t} \cos\left(z_{t+c} - z_t,\, g_t\right) \right]

where:

  • Rt=∑k=0∞γkrt+kR_t = \sum_{k=0}^\infty \gamma^k r_{t+k} is the cumulative discounted extrinsic reward from the environment.
  • VM(st)V_M(s_t) is the Manager's baseline value function.
  • AM=Rt−VM(st)A_M = R_t - V_M(s_t) is the Manager's advantage.

Intuition: If a sequence of actions produced high extrinsic reward (AM>0A_M > 0), the gradient pushes the Manager's emitted goal gtg_t to point directly along the realized displacement vector zt+c−ztz_{t+c} - z_t. The Manager learns to command directions that correlate with winning!

The Non-Hiding Principle

Earlier hierarchical RL architectures enforced "information hiding": the higher level hid sensory details from the lower level. This created artificial communication bottlenecks.

FuN adheres to the Non-Hiding Principle: both the Manager and Worker have full access to the shared perceptual representation ztz_t. Hierarchical abstraction is achieved not by concealing data, but through temporal dilation (the Manager only thinks every cc steps) and functional specialization (the Manager plans trajectories, the Worker actuates controls).

Worked numerical example

Let us trace a concrete macro-step of dilation horizon c=2c = 2 in a 2-dimensional latent representation space.

Setup

  • Initial latent state at t=0t=0: z0=[1.0,0.0]z_0 = [1.0, 0.0]
  • Manager directional goal: g0=[0.0,1.0]g_0 = [0.0, 1.0] (aiming along the positive yy-axis)
  • Dilation horizon: c=2c = 2

Step 1: Worker Execution (t=1t=1)

The Worker executes primitive action a0a_0, moving the environment to: z1=[1.0,0.6]z_1 = [1.0, 0.6]

The displacement vector is: Δz1=z1−z0=[1.0−1.0, 0.6−0.0]=[0.0, 0.6]\Delta z_1 = z_1 - z_0 = [1.0 - 1.0,\, 0.6 - 0.0] = [0.0,\, 0.6]

Compute cosine similarity with commanded goal g0=[0.0,1.0]g_0 = [0.0, 1.0]: Dot Product=0.0×0.0+0.6×1.0=0.60\text{Dot Product} = 0.0 \times 0.0 + 0.6 \times 1.0 = 0.60 ∥Δz1∥=0.02+0.62=0.60,∥g0∥=1.00\|\Delta z_1\| = \sqrt{0.0^2 + 0.6^2} = 0.60, \quad \|g_0\| = 1.00 cos⁡(Δz1,g0)=0.600.60×1.00=+1.0000\cos(\Delta z_1, g_0) = \frac{0.60}{0.60 \times 1.00} = +1.0000

The Worker's intrinsic reward at step 1 (with c=2c=2): r1I=12×cos⁡(Δz1,g0)=12×1.0000=0.5000r^I_1 = \frac{1}{2} \times \cos(\Delta z_1, g_0) = \frac{1}{2} \times 1.0000 = 0.5000

Step 2: Worker Execution (t=2t=2)

The Worker executes primitive action a1a_1, moving to: z2=[1.2,1.6]z_2 = [1.2, 1.6]

The total displacement from the macro-step origin z0z_0 is: Δz2=z2−z0=[1.2−1.0, 1.6−0.0]=[0.2, 1.6]\Delta z_2 = z_2 - z_0 = [1.2 - 1.0,\, 1.6 - 0.0] = [0.2,\, 1.6]

Compute cosine similarity: Dot Product=0.2×0.0+1.6×1.0=1.60\text{Dot Product} = 0.2 \times 0.0 + 1.6 \times 1.0 = 1.60 ∥Δz2∥=0.22+1.62=0.04+2.56=2.60≈1.61245\|\Delta z_2\| = \sqrt{0.2^2 + 1.6^2} = \sqrt{0.04 + 2.56} = \sqrt{2.60} \approx 1.61245 cos⁡(Δz2,g0)=1.601.61245×1.00≈0.992277≈0.9923\cos(\Delta z_2, g_0) = \frac{1.60}{1.61245 \times 1.00} \approx 0.992277 \approx 0.9923

The Worker's intrinsic reward contribution at step 2: r2I=12×0.992277≈0.4961r^I_2 = \frac{1}{2} \times 0.992277 \approx 0.4961

Step 3: Manager Transition Policy Gradient Update

At step t=2t=2, the environment emits an extrinsic reward R2=5.00R_2 = 5.00. The Manager's baseline value network estimated: VM(z0)=3.00V_M(z_0) = 3.00

The Manager advantage is: AM=R2−VM(z0)=5.00−3.00=+2.00A_M = R_2 - V_M(z_0) = 5.00 - 3.00 = +2.00

Because the advantage is positive (AM=+2.00A_M = +2.00), the Manager's Transition Policy Gradient updates g0g_0 to align even more strongly with the realized displacement vector [0.2,1.6][0.2, 1.6]. The Manager recognizes that directing the Worker into the upper-right quadrant produced high extrinsic returns!

Code

Below is a self-contained, type-hinted Python implementation of FeUdal Networks' directional goal calculation, Worker cosine-similarity intrinsic reward evaluation across dilation horizon cc, and Manager Transition Policy Gradient updates:

from dataclasses import dataclassimport mathfrom typing import List, Tuple

@dataclassclass FeUdalStepLog:    """Stores metrics for a single hierarchical step in FuN."""
    timestep: int    latent_state: List[float]    displacement: List[float]    cosine_sim: float    worker_intrinsic_reward: float

class FeUdalNetworkSimulator:    """Simulates the interaction between a dilated Manager and high-frequency Worker in FeUdal Networks."""
    def __init__(self, dilation_c: int = 2) -> None:        self.c = dilation_c
    def cosine_similarity(        self, vec_a: List[float], vec_b: List[float]    ) -> float:        """Computes cosine similarity: cos(a, b) = (a . b) / (||a|| * ||b||)."""        dot = sum(x * y for x, y in zip(vec_a, vec_b))        norm_a = math.sqrt(sum(x * x for x in vec_a))        norm_b = math.sqrt(sum(y * y for y in vec_b))        if norm_a == 0.0 or norm_b == 0.0:            return 0.0        return dot / (norm_a * norm_b)
    def evaluate_worker_step(        self,        current_z: List[float],        start_z: List[float],        goal_direction: List[float],    ) -> Tuple[List[float], float, float]:        """Calculates displacement, cosine alignment, and step intrinsic reward."""        disp = [curr - start for curr, start in zip(current_z, start_z)]        cos_sim = self.cosine_similarity(disp, goal_direction)        intrinsic_reward = (1.0 / self.c) * cos_sim        return disp, cos_sim, intrinsic_reward
    def compute_manager_transition_gradient(        self,        extrinsic_return_R: float,        manager_value_estimate: float,        total_disp: List[float],        goal: List[float],    ) -> Tuple[float, float]:        """Computes Manager advantage A_M = R - V_M and directional alignment."""        advantage = extrinsic_return_R - manager_value_estimate        alignment = self.cosine_similarity(total_disp, goal)        return advantage, alignment

if __name__ == "__main__":    sim = FeUdalNetworkSimulator(dilation_c=2)
    # 1. Worked Numerical Example Setup    # Macro-step origin z_0 and Manager unit goal g_0    z_0 = [1.0, 0.0]    g_0 = [0.0, 1.0]
    print("=== FeUdal Networks Macro-Step Simulation ===")    print(f"Origin Latent State z_0:      {z_0}")    print(f"Manager Directional Goal g_0: {g_0} (pointing along +y axis)")
    # 2. Worker Step 1: Moves to z_1 = [1.0, 0.6]    z_1 = [1.0, 0.6]    disp_1, cos_1, r_I_1 = sim.evaluate_worker_step(z_1, z_0, g_0)
    print("\n--- Worker Step 1 (t=1) ---")    print(f"Arrived State z_1:           {z_1}")    print(f"Displacement Delta z_1:      {disp_1}")    print(f"Cosine Similarity:           {cos_1:.4f}")    print(f"Worker Intrinsic Reward r_1: {r_I_1:.4f}")
    assert math.isclose(cos_1, 1.0000)    assert math.isclose(r_I_1, 0.5000)
    # 3. Worker Step 2: Moves to z_2 = [1.2, 1.6]    z_2 = [1.2, 1.6]    disp_2, cos_2, r_I_2 = sim.evaluate_worker_step(z_2, z_0, g_0)    norm_disp_2 = math.sqrt(disp_2[0] ** 2 + disp_2[1] ** 2)
    print("\n--- Worker Step 2 (t=2) ---")    print(f"Arrived State z_2:           {z_2}")    print(f"Displacement Delta z_2:      {disp_2}")    print(f"Displacement Norm:           {norm_disp_2:.5f}")    print(f"Cosine Similarity:           {cos_2:.4f}")    print(f"Worker Intrinsic Reward r_2: {r_I_2:.4f}")
    assert math.isclose(norm_disp_2, 1.61245, rel_tol=1e-4)    assert math.isclose(cos_2, 0.992277, rel_tol=1e-4)    assert math.isclose(r_I_2, 0.496138, rel_tol=1e-4)
    # 4. Manager Transition Policy Gradient Update    R_extrinsic = 5.00    V_manager = 3.00    advantage_M, final_alignment = (        sim.compute_manager_transition_gradient(            extrinsic_return_R=R_extrinsic,            manager_value_estimate=V_manager,            total_disp=disp_2,            goal=g_0,        )    )
    print("\n=== Manager Transition Policy Gradient Update ===")    print(f"Extrinsic Environment Return R: {R_extrinsic:.2f}")    print(f"Manager Value Baseline V_M:     {V_manager:.2f}")    print(f"Manager Advantage A_M:          {advantage_M:+.2f}")    print(f"Directional Alignment:          {final_alignment:.4f}")
    assert math.isclose(advantage_M, 2.00)    print("\nAll FeUdal Networks assertions verified successfully!")

Expected output:

=== FeUdal Networks Macro-Step Simulation ===Origin Latent State z_0:      [1.0, 0.0]Manager Directional Goal g_0: [0.0, 1.0] (pointing along +y axis)
--- Worker Step 1 (t=1) ---Arrived State z_1:           [1.0, 0.6]Displacement Delta z_1:      [0.0, 0.6]Cosine Similarity:           1.0000Worker Intrinsic Reward r_1: 0.5000
--- Worker Step 2 (t=2) ---Arrived State z_2:           [1.2, 1.6]Displacement Delta z_2:      [0.19999999999999996, 1.6]Displacement Norm:           1.61245Cosine Similarity:           0.9923Worker Intrinsic Reward r_2: 0.4961
=== Manager Transition Policy Gradient Update ===Extrinsic Environment Return R: 5.00Manager Value Baseline V_M:     3.00Manager Advantage A_M:          +2.00Directional Alignment:          0.9923
All FeUdal Networks assertions verified successfully!

Watch Out For

Goal Space Semantic Drift and Representation Collapse

The Trap: In FuN, the Manager emits directional goals in the perceptual representation space zt=fpercept(st)z_t = f_{\text{percept}}(s_t). If the weights of the perceptual encoder fperceptf_{\text{percept}} are trained jointly using policy gradients from the Manager or Worker, the geometry of the latent coordinate space drifts continuously. A goal vector gtg_t that meant "move toward the door" at step tt might mean "collide with the wall" ten gradient steps later because the underlying representation axis rotated.

The Symptom: The Worker's cosine-similarity intrinsic rewards fluctuate wildly; the agent experiences sudden catastrophic performance drops after achieving initial mastery, as previously learned motor behaviors no longer align with newly shifted goal coordinates.

The Fix:

  1. Stop Gradients into Perception: Explicitly block gradients from the Manager's Transition Policy Gradient and the Worker's policy gradient from backpropagating into the perceptual encoder fperceptf_{\text{percept}}.
  2. Self-Supervised Auxiliary Representation Learning: Train the perceptual representation ztz_t exclusively through auxiliary self-supervised tasks (such as future frame prediction, contrastive predictive coding, or autoencoding) to ensure a temporally stable, semantically grounded latent geometry.

The Quick Version

  • Two-Tier Temporal Abstraction: FuN decouples hierarchical control into a dilated Manager operating every cc steps (e.g., c=10c=10) and a high-frequency Worker operating at primitive step t=1t=1.
  • Directional Latent Subgoals: The Manager commands goals as directional unit vectors gt=dt/∥dt∥g_t = d_t / \|d_t\| in learned representation space zz, specifying which way to change the state rather than over-constraining the Worker with exact coordinates.
  • Cosine Intrinsic Reward: The Worker is motivated purely by internal cosine similarity between achieved state displacement Δz\Delta z and the commanded goal gtg_t, providing dense guidance in sparse-reward environments.
  • Transition Policy Gradients: The Manager is trained without backpropagating through unrolled Worker actions, directly reinforcing goals that aligned with state displacements yielding high extrinsic returns.