Radial Basis Functions in RL
Unlike grid discretization or tile coding that divide space with sharp binary edges, radial basis functions produce smooth, bell-shaped receptive fields. Each continuous state activates neighboring basis centers in proportion to its distance, yielding continuously differentiable value surfaces.
Why Does This Exist?
In continuous control problems—such as balancing an inverted pendulum, flying an autonomous drone, or navigating a robotic arm—state variables like angle, velocity, and torque vary continuously.
When applying linear function approximation to continuous state spaces, classical representations like state aggregation and Coarse Coding rely on binary membership features: a state either falls inside a discrete tile () or it does not ().
This discrete quantization introduces critical problems in gradient-based reinforcement learning:
- Discontinuous Step Artifacts: The approximated value function resembles a staircase of flat plateaus separated by sharp cliff boundaries. Crossing a tile boundary triggers sudden, jarring jumps in value predictions.
- Zero Gradients for Policy Optimization: In actor-critic and policy gradient methods, continuous action selection often relies on the spatial gradient of the value surface . On flat plateaus, the spatial gradient is identically zero (), providing zero directional guidance to the policy. At boundary edges, the gradient is non-differentiable or infinite.
Radial Basis Functions (RBFs) exist to provide smooth, continuous, and infinitely differentiable () feature representations.
Instead of hard binary partitions, RBFs deploy continuous bell-shaped receptive fields—most commonly Gaussian curves—centered at prototype coordinates across the state space. As an agent moves continuously through the environment, its feature activations shift smoothly between zero and one. This eliminates artificial boundary cliffs, enables analytical spatial gradients , and allows linear models to represent curved, organic value surfaces.
Think of It Like This
The Diffuse Theater Spotlights
Imagine walking across a dark theater stage illuminated by overhead spotlights aimed at three distinct floor positions: stage left (), center stage (), and stage right ().
These are not theatrical profile spots with sharp, knife-edge shutters that cast crisp geometric squares on the floor (which would be tile coding). Instead, each light is fitted with a soft diffusion frost lens. At the exact center of a spotlight's beam, the illumination is at peak brightness (). As you move away from the center, the brightness tapers off smoothly in all directions following a Gaussian bell curve.
Suppose you stand at coordinate (slightly to the right of center stage):
- Center Stage Spotlight () is only meters away: its beam strikes you with strong, direct illumination at intensity ().
- Stage Right Spotlight () is meters away: its peripheral halo casts a moderate illumination on you ().
- Stage Left Spotlight () is far away across the stage: its ambient fringe reaches you at a faint glow ().
Your overall visibility is the weighted sum of all three overlapping light beams. As you take a step, your lighting does not flicker, strobe, or jump abruptly across an invisible border. The light transitions seamlessly from one beam to the next.
Where the analogy stops: In optics, light beams combine strictly via positive physical superposition (you cannot project "negative light"). In reinforcement learning, RBF linear weights can be negative, allowing the learned value surface to represent deep negative penalty valleys and sharp reward peaks with equal fluidity.
How It Actually Works
Gaussian RBF Formulation, Bandwidth Tuning, and Semi-Gradient Updates
In an RBF network, a continuous state is transformed into a -dimensional continuous feature vector .
The Gaussian Basis Function
The standard Gaussian radial basis function is parameterized by a prototype center coordinate and a bandwidth standard deviation :
where:
- is the Euclidean distance between state and basis center .
- represents the degree of membership or activation.
- When , the exponent evaluates to zero, yielding maximal activation .
- As distance , activation decays exponentially toward zero.
For multi-dimensional spaces where dimensions have different physical units (e.g., position vs. velocity), each basis can use a diagonal covariance matrix :
Linear Value Approximation and Spatial Gradients
The state-value estimate is the linear combination of basis activations and learned weights :
Because the Gaussian exponential is smooth and infinitely differentiable, the spatial gradient with respect to the state coordinates is analytically available:
This analytical gradient allows continuous actor-critic algorithms to perform policy optimization directly via backpropagation into action parameters.
Semi-Gradient TD Update Rule
Under semi-gradient Temporal Difference Learning, observing transition produces scalar TD error:
Weights are updated in the direction of the basis gradient :
where is the step size. Crucially, every weight is updated in direct proportion to its active participation in the visited state.
Bandwidth Tuning: The Over-smoothing vs. Dead-Zone Dilemma
The bandwidth parameter governs the spatial width of generalization:
Receptive Field Envelopes Wide σ (Oversmoothing) Optimal σ (Smooth Overlap) Narrow σ (Dead Zones) ┌────────────────────────────┐ ┌────────────────────────────┐ ┌────────────────────────────┐ │ ╭──────╮ │ │ ╭──╮ ╭──╮ │ │ ││ ││ │ │ ╭───╯ ╰───╮ │ │ ╭─╯ ╰─╮ ╭─╯ ╰─╮ │ │ ╭╯╰╮ ╭╯╰╮ │ │ ────╯ ╰──── │ │ ──╯ ╰────╯ ╰── │ │ ──╯ ╰──────────╯ ╰── │ └────────────────────────────┘ └────────────────────────────┘ └────────────────────────────┘ Washes out sharp peaks Balanced generalization Zero transfer, gap artifacts- Wide (Excessive Overlap): Value updates generalize over wide regions of state space. However, localized reward peaks and fine decision boundaries are blurred into a featureless average.
- Narrow (Insufficient Overlap): Receptive fields barely touch. In regions between prototype centers, all basis activations drop to near zero (). These "dead zones" cause predictions to collapse to zero and freeze learning gradients.
- Calibrated Rule of Thumb: Set approximately equal to the spacing between adjacent centers () so that adjacent basis functions intersect at roughly to .
RBFs vs. Tile Coding
| Feature Attribute | Gaussian RBFs | Tile Coding |
|---|---|---|
| Surface Smoothness | continuously differentiable | Piecewise-constant step function |
| Boundary Artifacts | Zero boundary discontinuities | Step-edge artifacts at tile borders |
| Compute Complexity | transcendental exponentials | sparse integer additions |
| Sparsity | Dense (all bases active to some degree) | Strictly sparse (exactly active tiles) |
| High Dimensions | Suffers severe curse of dimensionality | Scalable via hashing and sparse tilings |
Worked numerical calculation
Consider an agent navigating a 1D continuous state space .
We configure a 3-basis Gaussian RBF network:
- Centers:
- Uniform bandwidth:
- Denominator:
- Initial weight vector:
- Step-size parameter:
We evaluate the continuous state and apply a semi-gradient update with observed target return .
Step 1: Compute Basis Feature Activations
- Basis 1 ():
- Basis 2 ():
- Basis 3 ():
The constructed feature vector is:
Step 2: Compute Predicted Value
Taking the inner product :
Step 3: Compute Prediction Error and Weight Updates
With target return :
Applying the update with :
- Weight 1 ():
- Weight 2 ():
- Weight 3 ():
Notice that Center 2 absorbs of the total update () because it is nearest to , while neighboring centers smoothly receive proportional fractional updates without boundary truncation.
Code
import mathfrom typing import List, Optional, Tuple
class GaussianRBFValueApproximator: """Linear value function approximator using continuous Gaussian Radial Basis Functions."""
def __init__( self, centers: List[float], sigma: float, initial_weights: Optional[List[float]] = None, ) -> None: """Initialize RBF network.
Args: centers: List of 1D center prototype coordinates c_i. sigma: Standard deviation / bandwidth parameter for all bases. initial_weights: Optional initial weight vector w. """ self.centers = centers self.sigma = sigma self.denom = 2.0 * (sigma ** 2) self.num_bases = len(centers) self.weights = ( initial_weights.copy() if initial_weights is not None else [0.0] * self.num_bases )
def compute_features(self, state: float) -> List[float]: """Compute continuous Gaussian RBF feature vector phi(s) in (0, 1]^d.
Args: state: Continuous 1D state coordinate s.
Returns: Vector of basis activations [phi_1(s), ..., phi_d(s)]. """ return [ math.exp(-((state - c) ** 2) / self.denom) for c in self.centers ]
def predict(self, state: float) -> float: """Evaluate estimated state value v_hat(s, w) = w^T phi(s).""" features = self.compute_features(state) return sum(w * phi for w, phi in zip(self.weights, features))
def update( self, state: float, target_value: float, alpha: float = 0.1, ) -> Tuple[float, List[float]]: """Perform a semi-gradient TD / Monte Carlo update step.
Args: state: Continuous state s where update target was observed. target_value: Sample return or TD target U_t. alpha: Step-size learning rate.
Returns: Tuple of (prediction_error delta, updated_weight_vector). """ features = self.compute_features(state) prediction = sum(w * phi for w, phi in zip(self.weights, features)) delta = target_value - prediction
for i in range(self.num_bases): self.weights[i] += alpha * delta * features[i]
return delta, self.weights.copy()
# Verification matching the worked numerical calculationrbf_model = GaussianRBFValueApproximator( centers=[0.0, 0.5, 1.0], sigma=0.3, initial_weights=[2.0, 5.0, 8.0],)
test_state = 0.60features = rbf_model.compute_features(test_state)
assert abs(features[0] - 0.1353) < 1e-4assert abs(features[1] - 0.9460) < 1e-4assert abs(features[2] - 0.4111) < 1e-4
initial_val = rbf_model.predict(test_state)assert abs(initial_val - 8.2894) < 1e-4
# Execute semi-gradient update with target return U = 10.0 and alpha = 0.1err, new_w = rbf_model.update(state=test_state, target_value=10.0, alpha=0.1)
assert abs(err - 1.7106) < 1e-4assert abs(new_w[0] - 2.0232) < 1e-4assert abs(new_w[1] - 5.1618) < 1e-4assert abs(new_w[2] - 8.0703) < 1e-4
print(f"RBF Features at s={test_state}: {[round(f, 4) for f in features]}")print(f"Predicted Value v_hat(s): {initial_val:.4f}")print(f"Prediction Error delta: {err:.4f}")print(f"Updated Weights w: {[round(w, 4) for w in new_w]}")
# -> RBF Features at s=0.6: [0.1353, 0.946, 0.4111]# -> Predicted Value v_hat(s): 8.2894# -> Prediction Error delta: 1.7106# -> Updated Weights w: [2.0232, 5.1618, 8.0703]Watch Out For
The Exponential Basis Scaling Curse and Unnormalized Dead Zones
Deploying RBF networks in multi-dimensional reinforcement learning introduces two severe failure modes:
1. Dimensionality Explosion and Computational Bottlenecks: Placing RBF centers on a regular grid across continuous dimensions requires basis functions. For an 8-dimensional robot arm with centers per dimension, the network requires centers. Unlike Tile Coding, where only a sparse handful () of tiles are active per step, RBFs are dense: every single update evaluates Euclidean distances and exponential transcendental functions, collapsing training throughput.
2. Unnormalized Dead Zones: If centers are placed too far apart relative to bandwidth , regions between prototypes have near-zero activations (). Value predictions in these dead zones collapse toward zero, and gradients vanish, trapping agents in unlearned regions.
Concrete Fix:
- Never use Cartesian grids in high dimensions: Sample RBF centers along actual agent trajectories, or use -means clustering on replay buffer observations to allocate centers exclusively to reachable regions of state space.
- Normalize activations via Normalized RBFs (Softmax Basis): Normalize basis activations so they form a partition of unity summing to : This eliminates dead zones entirely, ensuring every state receives a well-scaled value estimate and active learning gradient.
The Quick Version
- Radial Basis Functions provide continuous, infinitely differentiable Gaussian feature activations for linear reinforcement learning.
- By replacing sharp binary memberships with smooth bell curves, RBFs eliminate step-function artifacts and provide analytical spatial gradients .
- Bandwidth governs the generalization trade-off: wide washes out local value peaks, while narrow creates unlearned dead zones.
- While smooth and expressive in low dimensions, RBFs are computationally denser than tile coding ( transcendental exponentials vs sparse additions).