Feudal Networks (FuN)
A high-level Manager sets directional subgoals in a learned latent space every few steps, while a fast Worker learns primitive actions driven purely by cosine similarity to those goals.
Why Does This Exist?
In long-horizon reinforcement learning tasks with sparse rewards (such as the infamous Atari game Montezuma's Revenge), standard "flat" RL algorithms fail completely. When an agent must execute thousands of primitive actions before receiving a single scalar reward, credit assignment breaks down: the probability of discovering a rewarding trajectory through random exploratory motor noise is exponentially vanishing .
Hierarchical Reinforcement Learning (HRL) aims to solve this by decomposing complex behaviors into multi-level hierarchies. However, early HRL frameworks faced crippling bottlenecks:
- Handcrafted Subgoals: Traditional architectures required engineers to manually define subgoals (e.g., discrete room coordinates or specific inventory flags), destroying end-to-end learning.
- The Options Dilemma: Methods like the Options Framework required learning when to initiate and terminate options, suffering from option collapse where options degraded into single-step primitive actions.
- Information Hiding Bottlenecks: Classical Feudal RL (Dayan & Hinton, 1993) hid state details from higher levels, creating severe communication bandwidth limits.
FeUdal Networks (FuN), introduced by Alexander Vezhnevets et al. at DeepMind in 2017, modernized feudal reinforcement learning for deep neural networks. FuN separates control into two distinct recurrent tiers operating at different temporal resolutions: a Manager operating at a dilated macro-timescale , and a Worker operating at primitive timescale .
Crucially, FuN eliminated manual subgoal engineering by defining subgoals as directional unit vectors in a learned continuous latent space. The Worker is motivated purely by cosine-similarity intrinsic rewards, while the Manager is trained using a novel Transition Policy Gradient that eliminates the need to backpropagate gradients through the Worker's motor decisions.
Think of It Like This
The Medieval Feudal Hierarchy
Imagine the administrative structure of a medieval agrarian kingdom:
The Feudal Lord (Manager) sits in the high castle tower observing kingdom-wide crop yields, seasonal rainfall, and long-range logistics. The Lord summons the local Baron (Worker) and issues an executive directional decree:
"Expand our agricultural territory toward the eastern valley."
The Lord does not know how to dig irrigation ditches, forge iron hoes, or swing scythes, and does not micromanage which serf plows which furrow. The Lord simply specifies the high-level directional objective in strategic coordinates.
The Baron (Worker) lives on the ground with the farmers. The Baron receives personal wealth and prestige (Intrinsic Reward ) proportionally to how effectively the workforce clears land in that commanded eastward direction (). The Baron does not need to worry about kingdom-wide trade treaties or royal treasury reserves (Extrinsic Reward )—they focus purely on solving local physical obstacles (boulders, rivers, mud) to satisfy the Lord's directional vector.
At harvest time, if the kingdom flourishes ( is high), the Lord receives royal credit. The Lord updates their strategic playbook (Transition Policy Gradient) to issue similar successful territorial commands in future seasons.
Where the analogy stops: In human feudalism, lords and barons communicate through spoken language that can be misunderstood, and information is often hidden. In FuN, both Manager and Worker share access to the same learned perceptual state (the Non-Hiding Principle), ensuring hierarchical abstraction without communication bottlenecks.
How It Actually Works
Architectural Decoupling: Manager and Worker
FuN takes raw environment observations (such as visual image frames) and passes them through a shared convolutional perception backbone to produce a low-dimensional latent state:
The architecture then splits into two specialized recurrent modules:
Observation s_t ──> [ Perception Backbone ] ──> Latent Representation z_t │ ┌───────────────────────────────────────────┴───────────────────────────────────────────┐ ▼ ▼┌─────────────────────────────────────────┐ ┌─────────────────────────────────────────┐│ Dilated Manager RNN (Macro-step c) │ │ Fast Worker RNN (Micro-step t = 1) ││ Ticks every c steps (e.g., c = 10) │ │ Ticks every single environment step ││ Outputs Directional Goal g_t ∈ ℝ^d │ │ Takes z_t and Goal Embedding │└─────────────────────────────────────────┘ └─────────────────────────────────────────┘ │ │ ▼ ▼ Directional Goal g_t Primitive Motor Action a_t │ │ └───────────────────────────┬───────────────────────────────────────────────────────────┘ ▼ ┌─────────────────────────────────────────┐ │ Latent Cosine Alignment Module │ │ Measures: cos(z_t - z_{t-i}, g_{t-i}) │ │ Yields: Worker Intrinsic Reward r^I_t │ └─────────────────────────────────────────┘1. The Dilated Manager (Macro-timescale )
The Manager operates at a coarse temporal resolution, ticking once every steps (typically ). To maintain temporal memory across hundreds of steps without vanishing gradients, the Manager uses a Dilated LSTM.
At macro-step , the Manager outputs a continuous latent goal vector , normalized to a unit vector:
Key Insight: is not an absolute target coordinate in state space. It is a directional displacement vector, commanding: "Over the next steps, shift the environment's latent state in direction ."
2. The Fast Worker (Primitive timescale )
The Worker operates at every single environment step . It receives the perceptual state and an internal goal embedding . It produces a probability distribution over primitive motor actions:
The Worker handles low-level agility, collision avoidance, and motor reflexes.
Worker Intrinsic Reward via Cosine Similarity
The Worker is completely shielded from external environment rewards . Instead, it is motivated purely by an internal cosine-similarity intrinsic reward .
At each step , the intrinsic reward measures how closely the actual state displacement achieved over the past steps aligns with the Manager's commanded directional goal:
where the cosine similarity between two vectors and is:
If the Worker's primitive actions move the agent in the exact direction commanded by the Manager, , maximizing intrinsic payoff. If the Worker moves perpendicular to the goal, reward is ; if it moves in the opposite direction, reward is .
The Worker maximizes using standard policy gradient algorithms (such as A3C or PPO).
Transition Policy Gradients for the Manager
Training a hierarchical manager usually creates a mathematical dilemma: how can the Manager receive gradients without backpropagating through hundreds of unrolled Worker actions?
Vezhnevets et al. introduced Transition Policy Gradients by exploiting the structure of the goal space. The Worker is actively trained to align state displacements with goal . Therefore, from the Manager's perspective, the probability of the environment transitioning from to can be modeled as:
Applying the policy gradient theorem to this transition distribution yields the Manager's update rule:
where:
- is the cumulative discounted extrinsic reward from the environment.
- is the Manager's baseline value function.
- is the Manager's advantage.
Intuition: If a sequence of actions produced high extrinsic reward (), the gradient pushes the Manager's emitted goal to point directly along the realized displacement vector . The Manager learns to command directions that correlate with winning!
The Non-Hiding Principle
Earlier hierarchical RL architectures enforced "information hiding": the higher level hid sensory details from the lower level. This created artificial communication bottlenecks.
FuN adheres to the Non-Hiding Principle: both the Manager and Worker have full access to the shared perceptual representation . Hierarchical abstraction is achieved not by concealing data, but through temporal dilation (the Manager only thinks every steps) and functional specialization (the Manager plans trajectories, the Worker actuates controls).
Worked numerical example
Let us trace a concrete macro-step of dilation horizon in a 2-dimensional latent representation space.
Setup
- Initial latent state at :
- Manager directional goal: (aiming along the positive -axis)
- Dilation horizon:
Step 1: Worker Execution ()
The Worker executes primitive action , moving the environment to:
The displacement vector is:
Compute cosine similarity with commanded goal :
The Worker's intrinsic reward at step 1 (with ):
Step 2: Worker Execution ()
The Worker executes primitive action , moving to:
The total displacement from the macro-step origin is:
Compute cosine similarity:
The Worker's intrinsic reward contribution at step 2:
Step 3: Manager Transition Policy Gradient Update
At step , the environment emits an extrinsic reward . The Manager's baseline value network estimated:
The Manager advantage is:
Because the advantage is positive (), the Manager's Transition Policy Gradient updates to align even more strongly with the realized displacement vector . The Manager recognizes that directing the Worker into the upper-right quadrant produced high extrinsic returns!
Code
Below is a self-contained, type-hinted Python implementation of FeUdal Networks' directional goal calculation, Worker cosine-similarity intrinsic reward evaluation across dilation horizon , and Manager Transition Policy Gradient updates:
from dataclasses import dataclassimport mathfrom typing import List, Tuple
@dataclassclass FeUdalStepLog: """Stores metrics for a single hierarchical step in FuN."""
timestep: int latent_state: List[float] displacement: List[float] cosine_sim: float worker_intrinsic_reward: float
class FeUdalNetworkSimulator: """Simulates the interaction between a dilated Manager and high-frequency Worker in FeUdal Networks."""
def __init__(self, dilation_c: int = 2) -> None: self.c = dilation_c
def cosine_similarity( self, vec_a: List[float], vec_b: List[float] ) -> float: """Computes cosine similarity: cos(a, b) = (a . b) / (||a|| * ||b||).""" dot = sum(x * y for x, y in zip(vec_a, vec_b)) norm_a = math.sqrt(sum(x * x for x in vec_a)) norm_b = math.sqrt(sum(y * y for y in vec_b)) if norm_a == 0.0 or norm_b == 0.0: return 0.0 return dot / (norm_a * norm_b)
def evaluate_worker_step( self, current_z: List[float], start_z: List[float], goal_direction: List[float], ) -> Tuple[List[float], float, float]: """Calculates displacement, cosine alignment, and step intrinsic reward.""" disp = [curr - start for curr, start in zip(current_z, start_z)] cos_sim = self.cosine_similarity(disp, goal_direction) intrinsic_reward = (1.0 / self.c) * cos_sim return disp, cos_sim, intrinsic_reward
def compute_manager_transition_gradient( self, extrinsic_return_R: float, manager_value_estimate: float, total_disp: List[float], goal: List[float], ) -> Tuple[float, float]: """Computes Manager advantage A_M = R - V_M and directional alignment.""" advantage = extrinsic_return_R - manager_value_estimate alignment = self.cosine_similarity(total_disp, goal) return advantage, alignment
if __name__ == "__main__": sim = FeUdalNetworkSimulator(dilation_c=2)
# 1. Worked Numerical Example Setup # Macro-step origin z_0 and Manager unit goal g_0 z_0 = [1.0, 0.0] g_0 = [0.0, 1.0]
print("=== FeUdal Networks Macro-Step Simulation ===") print(f"Origin Latent State z_0: {z_0}") print(f"Manager Directional Goal g_0: {g_0} (pointing along +y axis)")
# 2. Worker Step 1: Moves to z_1 = [1.0, 0.6] z_1 = [1.0, 0.6] disp_1, cos_1, r_I_1 = sim.evaluate_worker_step(z_1, z_0, g_0)
print("\n--- Worker Step 1 (t=1) ---") print(f"Arrived State z_1: {z_1}") print(f"Displacement Delta z_1: {disp_1}") print(f"Cosine Similarity: {cos_1:.4f}") print(f"Worker Intrinsic Reward r_1: {r_I_1:.4f}")
assert math.isclose(cos_1, 1.0000) assert math.isclose(r_I_1, 0.5000)
# 3. Worker Step 2: Moves to z_2 = [1.2, 1.6] z_2 = [1.2, 1.6] disp_2, cos_2, r_I_2 = sim.evaluate_worker_step(z_2, z_0, g_0) norm_disp_2 = math.sqrt(disp_2[0] ** 2 + disp_2[1] ** 2)
print("\n--- Worker Step 2 (t=2) ---") print(f"Arrived State z_2: {z_2}") print(f"Displacement Delta z_2: {disp_2}") print(f"Displacement Norm: {norm_disp_2:.5f}") print(f"Cosine Similarity: {cos_2:.4f}") print(f"Worker Intrinsic Reward r_2: {r_I_2:.4f}")
assert math.isclose(norm_disp_2, 1.61245, rel_tol=1e-4) assert math.isclose(cos_2, 0.992277, rel_tol=1e-4) assert math.isclose(r_I_2, 0.496138, rel_tol=1e-4)
# 4. Manager Transition Policy Gradient Update R_extrinsic = 5.00 V_manager = 3.00 advantage_M, final_alignment = ( sim.compute_manager_transition_gradient( extrinsic_return_R=R_extrinsic, manager_value_estimate=V_manager, total_disp=disp_2, goal=g_0, ) )
print("\n=== Manager Transition Policy Gradient Update ===") print(f"Extrinsic Environment Return R: {R_extrinsic:.2f}") print(f"Manager Value Baseline V_M: {V_manager:.2f}") print(f"Manager Advantage A_M: {advantage_M:+.2f}") print(f"Directional Alignment: {final_alignment:.4f}")
assert math.isclose(advantage_M, 2.00) print("\nAll FeUdal Networks assertions verified successfully!")Expected output:
=== FeUdal Networks Macro-Step Simulation ===Origin Latent State z_0: [1.0, 0.0]Manager Directional Goal g_0: [0.0, 1.0] (pointing along +y axis)
--- Worker Step 1 (t=1) ---Arrived State z_1: [1.0, 0.6]Displacement Delta z_1: [0.0, 0.6]Cosine Similarity: 1.0000Worker Intrinsic Reward r_1: 0.5000
--- Worker Step 2 (t=2) ---Arrived State z_2: [1.2, 1.6]Displacement Delta z_2: [0.19999999999999996, 1.6]Displacement Norm: 1.61245Cosine Similarity: 0.9923Worker Intrinsic Reward r_2: 0.4961
=== Manager Transition Policy Gradient Update ===Extrinsic Environment Return R: 5.00Manager Value Baseline V_M: 3.00Manager Advantage A_M: +2.00Directional Alignment: 0.9923
All FeUdal Networks assertions verified successfully!Watch Out For
Goal Space Semantic Drift and Representation Collapse
The Trap: In FuN, the Manager emits directional goals in the perceptual representation space . If the weights of the perceptual encoder are trained jointly using policy gradients from the Manager or Worker, the geometry of the latent coordinate space drifts continuously. A goal vector that meant "move toward the door" at step might mean "collide with the wall" ten gradient steps later because the underlying representation axis rotated.
The Symptom: The Worker's cosine-similarity intrinsic rewards fluctuate wildly; the agent experiences sudden catastrophic performance drops after achieving initial mastery, as previously learned motor behaviors no longer align with newly shifted goal coordinates.
The Fix:
- Stop Gradients into Perception: Explicitly block gradients from the Manager's Transition Policy Gradient and the Worker's policy gradient from backpropagating into the perceptual encoder .
- Self-Supervised Auxiliary Representation Learning: Train the perceptual representation exclusively through auxiliary self-supervised tasks (such as future frame prediction, contrastive predictive coding, or autoencoding) to ensure a temporally stable, semantically grounded latent geometry.
The Quick Version
- Two-Tier Temporal Abstraction: FuN decouples hierarchical control into a dilated Manager operating every steps (e.g., ) and a high-frequency Worker operating at primitive step .
- Directional Latent Subgoals: The Manager commands goals as directional unit vectors in learned representation space , specifying which way to change the state rather than over-constraining the Worker with exact coordinates.
- Cosine Intrinsic Reward: The Worker is motivated purely by internal cosine similarity between achieved state displacement and the commanded goal , providing dense guidance in sparse-reward environments.
- Transition Policy Gradients: The Manager is trained without backpropagating through unrolled Worker actions, directly reinforcing goals that aligned with state displacements yielding high extrinsic returns.