Adversarial Inverse RL (AIRL)
Instead of just copying an expert's motions, AIRL uncovers the true underlying goal so the agent can still succeed when the environment changes.
Why Does This Exist?
Imitation learning methods like Generative Adversarial Imitation Learning (GAIL) train an agent to mimic expert behavior by matching state-action occupancy measures . While GAIL efficiently trains an imitating policy, its learned discriminator is fundamentally entangled with the training environment's transition dynamics .
If the environment changes—such as when a robot suffers a jammed motor, moves to a surface with different friction, or encounters newly placed obstacles ()—GAIL's discriminator outputs meaningless values. Because GAIL learned how to move rather than what the goal is, the policy fails completely.
To enable true policy transfer across changing physical dynamics, the agent must recover the ground-truth reward function through Inverse Reinforcement Learning (IRL). However, classical Maximum Entropy IRL (Ziebart et al., 2008) requires running an entire forward reinforcement learning loop at every optimization step, making it computationally prohibitive for complex continuous control tasks.
Adversarial Inverse Reinforcement Learning (AIRL) (Fu, Luo, & Levine, 2018) solves both problems simultaneously:
- It preserves the sample efficiency of adversarial training (updating the policy with single-step transitions).
- It mathematically decomposes the discriminator to disentangle the true state reward from transition-dependent potential shaping .
- The recovered reward function is provably invariant to environment dynamics, allowing an agent to re-plan and succeed in radically modified environments.
Think of It Like This
Teaching an agility course to a dog: Motion mimicry vs. Goal understanding
Imagine training an agility dog. A champion handler demonstrates an agility course: sprinting through a tunnel, leaping over a red hurdle, and crossing the finish line where a bowl of food is placed.
-
The GAIL Approach (Motion Mimicry): The student dog memorizes the exact running speed, leap timing, and paw placements from the demonstration.
As long as the hurdles remain in their exact original spots, the dog performs well. But if you move the red hurdle three meters to the left, the dog leaps into empty air, crashes into the displaced obstacle, and gets disqualified. It mimicked the physical motion without understanding the objective.
-
The AIRL Approach (Reward Disentanglement): The student dog observes the demonstration and separates the task into two components:
- Shaping Potential (): Jumping over hurdles was just the temporary physical maneuvering needed to navigate the specific arena layout.
- Ground-Truth Reward (): The true prize is reaching the finish line to eat the food.
When placed in a new arena with rearranged hurdles, the dog does not blindly replay the old paw placements. It plans a new path around the shifted hurdles and sprints straight to the finish line.
Where the analogy stops: A dog has biological senses to smell food and perceive rewards directly. In inverse reinforcement learning, the environment provides zero reward signals to the agent during demonstrations. AIRL must isolate the true reward entirely through the mathematical structure of its adversarial discriminator.
How It Actually Works
The Structured Discriminator and Reward Disentanglement
In Maximum Entropy Inverse Reinforcement Learning, the optimal policy takes the energy-based form .
AIRL incorporates this formulation directly into a GAN discriminator by structuring to contrast a learned advantage function against the generator policy probability :
1. Reward Disentanglement Parameterization
According to policy invariance under reward shaping (Ng, Harada, & Russell, 1999), two reward functions share the same optimal policy if and only if they differ by a potential difference .
AIRL explicitly parameterizes the advantage function to separate the true reward from the shaping potential:
- : A neural network that estimates the ground-truth reward, restricted to depend only on state .
- : A neural network that estimates the state potential function, absorbing transition dynamics and value baseline differences.
2. The Theoretical Guarantee (Fu et al., 2018)
When is parameterized as a state-only function , the discriminator's global optimum satisfies:
where is the true environment reward and is a constant. The learned reward is disentangled from transition dynamics , meaning it can be directly exported to an environment with different physics or obstacle layouts.
3. Adversarial Training Loop
Training alternates between updating the discriminator and the generator policy:
- Discriminator Update: Train using binary cross-entropy on expert transitions vs. policy transitions:
- Policy Update: Train policy via policy gradients (e.g. TRPO or PPO) using the surrogate reward:
Worked numerical example
Let us trace a single transition update under AIRL.
Transition Setup:
- Current state , action , and next state .
- Learner policy probability: .
- Learned state reward network: .
- Learned potential function evaluations: , .
- Discount factor: .
Step 1: Compute Shaped Advantage Function
Step 2: Compute Discriminator Output
Evaluate the exponential advantage:
Evaluating the structured discriminator:
Because , the discriminator strongly suspects this transition came from an expert.
Step 3: Compute Policy Surrogate Reward
Step 4: Dynamics Transfer Demonstration
Suppose the robot transitions to a modified environment with higher friction, causing action from state to transition to state instead of , where .
- The shaping term shifts: .
- The advantage shifts: .
- Crucially, the recovered reward function remains identical.
The agent transfers directly into environment and re-optimizes a policy that accommodates the new friction.
Code
import mathfrom typing import Dict, Tuple
class AdversarialInverseRLSimulator: """Demonstrates Adversarial Inverse Reinforcement Learning (AIRL; Fu et al., 2018):
1. Structured Discriminator: D = exp(f) / [exp(f) + pi(a | s)] 2. Advantage Decomposition: f = g_theta(s) + gamma * h_phi(s') - h_phi(s) 3. Surrogate Reward: r_hat = f - ln pi(a | s) 4. Reward Transfer across modified transition dynamics """
def __init__(self, gamma: float = 0.90) -> None: self.gamma = gamma
def compute_shaped_advantage( self, g_reward: float, h_current: float, h_next: float, ) -> float: """Calculates f(s, a, s') = g(s) + gamma * h(s') - h(s).""" return round(g_reward + self.gamma * h_next - h_current, 4)
@staticmethod def compute_discriminator_output( advantage_f: float, policy_prob: float, ) -> Tuple[float, float]: """Calculates D(s, a, s') = exp(f) / [exp(f) + pi(a | s)].""" exp_f = math.exp(advantage_f) d_val = exp_f / (exp_f + policy_prob) return round(exp_f, 4), round(d_val, 4)
@staticmethod def compute_surrogate_reward( advantage_f: float, policy_prob: float, ) -> float: """Calculates policy update reward: r_hat = f - ln pi(a | s).""" log_pi = math.log(policy_prob) return round(advantage_f - log_pi, 4)
# Initialize simulator with worked numerical example parametersairl = AdversarialInverseRLSimulator(gamma=0.90)
# 1. Primary Environment P_1 Evaluationg_state = 2.0 # Disentangled ground-truth reward g_theta(s)h_curr = 1.0 # State potential h_phi(s)h_next_p1 = 1.5 # State potential in environment P_1: h_phi(s')pi_a_s = 0.20 # Policy probability pi(a | s)
f_p1 = airl.compute_shaped_advantage(g_state, h_curr, h_next_p1)exp_f_p1, d_p1 = airl.compute_discriminator_output(f_p1, pi_a_s)r_surrogate_p1 = airl.compute_surrogate_reward(f_p1, pi_a_s)
print("=== Environment P_1 Evaluation ===")print(f"Shaped Advantage f(s, a, s'): {f_p1:.2f}")# -> Shaped Advantage f(s, a, s'): 2.35print(f"exp(f): {exp_f_p1:.4f}")# -> exp(f): 10.4856print(f"Discriminator D: {d_p1:.4f}")# -> Discriminator D: 0.9813print(f"Surrogate Reward r_hat: {r_surrogate_p1:.4f}")# -> Surrogate Reward r_hat: 3.9594print(f"Disentangled State Reward g_theta(s): {g_state:.2f}")# -> Disentangled State Reward g_theta(s): 2.00
# 2. Transfer to Modified Dynamics P_2 (Altered Friction / Obstacles)h_next_p2 = 0.80 # Next state potential shifts under new physics: h_phi(s'')f_p2 = airl.compute_shaped_advantage(g_state, h_curr, h_next_p2)
print("\n=== Environment P_2 (Modified Dynamics) Transfer ===")print( f"New Shaped Advantage f(s, a, s''): {f_p2:.2f} (absorbed by potential h)")# -> New Shaped Advantage f(s, a, s''): 1.72 (absorbed by potential h)print(f"Exported Ground-Truth Reward g_theta(s): {g_state:.2f} (UNCHANGED!)")# -> Exported Ground-Truth Reward g_theta(s): 2.00 (UNCHANGED!)
# Verification assertionsassert f_p1 == 2.35assert exp_f_p1 == 10.4856assert d_p1 == 0.9813assert r_surrogate_p1 == 3.9594assert g_state == 2.00assert f_p2 == 1.72Watch Out For
Action-Dependent Reward Ambiguity and Entanglement
A frequent mistake when extending AIRL is allowing the reward network to condition on actions: .
While action-dependent rewards seem more expressive, conditioning on action breaks the uniqueness guarantees of potential-based reward shaping. If depends on , the potential difference can no longer uniquely absorb the environment's transition dynamics . The learned reward function becomes mathematically entangled with the specific transition probabilities of the training environment.
When an action-dependent reward is transferred to an environment with modified dynamics (such as a robotic arm with an altered payload or altered joint damping), the agent learns erratic, divergent policies that exploit the obsolete shaping artifacts.
The Fix: Always restrict to be a state-only function . Fu et al. (Theorem 5.1) prove that when is state-only, the learned reward function is guaranteed to recover the true ground-truth reward up to an additive constant, ensuring robust transfer across arbitrary dynamic shifts.
The Quick Version
- AIRL recovers transferable ground-truth reward functions by structuring the adversarial discriminator: .
- The advantage function decomposes into disentangled state reward and potential shaping: .
- Unlike GAIL (which mimics motion trajectories that collapse when physics change), AIRL recovers the underlying goal, enabling seamless transfer to altered dynamics environments.
- To guarantee mathematical reward disentanglement and prevent dynamics leakage, must be strictly constrained to be a state-only function .