Skip to content
AI360Xpert
Beta

Calculus and Optimization

Calculus provides the mathematical language of instantaneous change, while optimization uses those rates to systematically steer parameters toward the lowest possible error.

Taylor series quadratic model guides parameter updates through loss landscapes toward minima.
Taylor series quadratic model guides parameter updates through loss landscapes toward minima.

Why Does This Exist?

Machine learning models do not learn through magical intuition; they solve numerical optimization problems. A neural network with dd weights is a parameterized function fθf_{\mathbf{\theta}}, and training seeks a weight vector θ∗∈Rd\mathbf{\theta}^* \in \mathbb{R}^d that minimizes an empirical loss function L(θ)\mathcal{L}(\mathbf{\theta}).

In modern models, parameter dimension dd ranges from thousands to hundreds of billions. Evaluating the loss at every possible point in parameter space is physically impossible — an exhaustive grid search over just 50 parameters with 10 values each requires 105010^{50} function evaluations. Random guessing is equally hopeless in high-dimensional space.

Calculus makes high-dimensional optimization tractable by replacing global search with local sensitivity analysis. Instead of wondering what the loss function looks like across the entire universe of weights, calculus asks a local question: if we alter each parameter by an infinitesimal amount, in what direction does the loss decline fastest, and how does the slope curve? Calculus computes the gradient and curvature, and optimization builds the iterative algorithms that ride those sensitivities to the lowest error.

For the foundations of directional derivatives and slope vectors, see our guides on derivatives and partial derivatives and gradients.

Think of It Like This

A blindfolded skier feeling slope and curvature through ski poles

Imagine standing on a snow-covered mountain in thick fog. You cannot see the valley floor, the surrounding peaks, or the distant ski lodge. Your goal is to descend to the lowest point on the mountain as quickly and safely as possible.

Your ski boots feel the tilt of the snow directly underneath you. That tilt is the first derivative — the gradient. It tells you which compass direction points downhill and how steep the immediate decline is. If you only use your boots, you take small steps in the steepest downhill direction.

Now you extend two ski poles and tap the snow five feet in front of you, behind you, and to your sides. By sensing how the slope changes between your boots and the pole tips, you detect curvature — the second derivative, represented by the Hessian. The poles tell you whether the slope is flattening out into a safe bowl, plummeting over a vertical cliff, or twisting into a treacherous mountain pass (a saddle point).

Optimization is your downhill strategy: calculus gathers the slope and curvature readings, and optimization decides how long your stride should be, when to accelerate, and when to brake to avoid crashing.

The analogy stops because mountains exist in 3 physical dimensions where gravity pulls downward automatically, whereas neural networks operate in billions of mathematical dimensions where optimization must compute its own synthetic gravity at every step.

How It Actually Works

Multivariable Taylor Approximations and Critical Points

At the heart of optimization is the multivariable Taylor series expansion. For any twice-differentiable loss function L:Rd→R\mathcal{L}: \mathbb{R}^d \to \mathbb{R}, we can approximate the loss in a local neighborhood around parameter vector θ0\mathbf{\theta}_0 by perturbing it with a small displacement vector Δθ\Delta \mathbf{\theta}:

L(θ0+Δθ)≈L(θ0)+∇L(θ0)TΔθ+12ΔθTH(θ0)Δθ\mathcal{L}(\mathbf{\theta}_0 + \Delta \mathbf{\theta}) \approx \mathcal{L}(\mathbf{\theta}_0) + \nabla \mathcal{L}(\mathbf{\theta}_0)^T \Delta \mathbf{\theta} + \frac{1}{2} \Delta \mathbf{\theta}^T \mathbf{H}(\mathbf{\theta}_0) \Delta \mathbf{\theta}

where ∇L(θ0)∈Rd\nabla \mathcal{L}(\mathbf{\theta}_0) \in \mathbb{R}^d is the gradient vector of first-order partial derivatives:

∇L(θ0)=[∂L∂θ1,∂L∂θ2,…,∂L∂θd]T\nabla \mathcal{L}(\mathbf{\theta}_0) = \left[ \frac{\partial \mathcal{L}}{\partial \theta_1}, \frac{\partial \mathcal{L}}{\partial \theta_2}, \dots, \frac{\partial \mathcal{L}}{\partial \theta_d} \right]^T

and H(θ0)∈Rd×d\mathbf{H}(\mathbf{\theta}_0) \in \mathbb{R}^{d \times d} is the symmetric Hessian matrix of second-order partial derivatives:

Hij=∂2L∂θi∂θjH_{ij} = \frac{\partial^2 \mathcal{L}}{\partial \theta_i \partial \theta_j}

The gradient provides the tangent hyperplane (local linear trend), while the Hessian provides the osculating paraboloid (local curvature).

First-Order Optimization: Gradient Descent

If we truncate the Taylor expansion to first order and add an L2L_2 step-size penalty 12η∥Δθ∥2\frac{1}{2\eta} \|\Delta \mathbf{\theta}\|^2 to keep the displacement within the valid local trust region, we minimize:

min⁡Δθ(∇L(θ0)TΔθ+12η∥Δθ∥2)\min_{\Delta \mathbf{\theta}} \left( \nabla \mathcal{L}(\mathbf{\theta}_0)^T \Delta \mathbf{\theta} + \frac{1}{2\eta} \|\Delta \mathbf{\theta}\|^2 \right)

Differentiating with respect to Δθ\Delta \mathbf{\theta} and setting to zero yields the foundational Gradient Descent update rule:

Δθ=−η∇L(θ0)  ⟹  θt+1=θt−η∇L(θt)\Delta \mathbf{\theta} = -\eta \nabla \mathcal{L}(\mathbf{\theta}_0) \implies \mathbf{\theta}_{t+1} = \mathbf{\theta}_t - \eta \nabla \mathcal{L}(\mathbf{\theta}_t)

where η>0\eta > 0 is the learning rate.

Second-Order Optimization: Newton's Method

If we include the quadratic term and differentiate the full second-order Taylor approximation with respect to Δθ\Delta \mathbf{\theta}:

∇Δθ(L(θ0)+∇L(θ0)TΔθ+12ΔθTHΔθ)=∇L(θ0)+HΔθ=0\nabla_{\Delta \mathbf{\theta}} \left( \mathcal{L}(\mathbf{\theta}_0) + \nabla \mathcal{L}(\mathbf{\theta}_0)^T \Delta \mathbf{\theta} + \frac{1}{2} \Delta \mathbf{\theta}^T \mathbf{H} \Delta \mathbf{\theta} \right) = \nabla \mathcal{L}(\mathbf{\theta}_0) + \mathbf{H} \Delta \mathbf{\theta} = \mathbf{0}

Solving for Δθ\Delta \mathbf{\theta} gives the Newton-Raphson update:

Δθ=−H(θ0)−1∇L(θ0)  ⟹  θt+1=θt−H(θt)−1∇L(θt)\Delta \mathbf{\theta} = -\mathbf{H}(\mathbf{\theta}_0)^{-1} \nabla \mathcal{L}(\mathbf{\theta}_0) \implies \mathbf{\theta}_{t+1} = \mathbf{\theta}_t - \mathbf{H}(\mathbf{\theta}_t)^{-1} \nabla \mathcal{L}(\mathbf{\theta}_t)

Newton's method rescales the step by the inverse curvature along every eigen-axis, jumping directly to the minimum of a quadratic bowl in a single step without needing a manual learning rate.

Critical Point Classification

A point where the gradient vanishes (∇L(θ∗)=0\nabla \mathcal{L}(\mathbf{\theta}^*) = \mathbf{0}) is a stationary (critical) point. The eigenvalues λ1,λ2,…,λd\lambda_1, \lambda_2, \dots, \lambda_d of the Hessian matrix H(θ∗)\mathbf{H}(\mathbf{\theta}^*) determine its geometric identity:

  1. Local Minimum: H≻0\mathbf{H} \succ 0 (all λi>0\lambda_i > 0, positive definite). The surface curves upward in every direction.
  2. Local Maximum: H≺0\mathbf{H} \prec 0 (all λi<0\lambda_i < 0, negative definite). The surface curves downward in every direction.
  3. Saddle Point: H\mathbf{H} is indefinite (at least one λi>0\lambda_i > 0 and at least one λj<0\lambda_j < 0). The surface rises along some axes and falls along others.

Worked Example

Consider minimizing the anisotropic quadratic loss function:

L(θ1,θ2)=2θ12+θ22−4θ1−2θ2+5\mathcal{L}(\theta_1, \theta_2) = 2\theta_1^2 + \theta_2^2 - 4\theta_1 - 2\theta_2 + 5

Starting at the origin θ0=[0,0]T\mathbf{\theta}_0 = [0, 0]^T:

  1. Compute first-order partial derivatives:

    ∂L∂θ1=4θ1−4,∂L∂θ2=2θ2−2\frac{\partial \mathcal{L}}{\partial \theta_1} = 4\theta_1 - 4, \quad \frac{\partial \mathcal{L}}{\partial \theta_2} = 2\theta_2 - 2

    At θ0=[0,0]T\mathbf{\theta}_0 = [0, 0]^T, the gradient is:

    ∇L(0,0)=[4(0)−42(0)−2]=[−4−2]\nabla \mathcal{L}(0, 0) = \begin{bmatrix} 4(0) - 4 \\ 2(0) - 2 \end{bmatrix} = \begin{bmatrix} -4 \\ -2 \end{bmatrix}
  2. Compute the second-order Hessian matrix:

    H=[∂2L∂θ12∂2L∂θ1∂θ2∂2L∂θ2∂θ1∂2L∂θ22]=[4002]\mathbf{H} = \begin{bmatrix} \frac{\partial^2 \mathcal{L}}{\partial \theta_1^2} & \frac{\partial^2 \mathcal{L}}{\partial \theta_1 \partial \theta_2} \\ \frac{\partial^2 \mathcal{L}}{\partial \theta_2 \partial \theta_1} & \frac{\partial^2 \mathcal{L}}{\partial \theta_2^2} \end{bmatrix} = \begin{bmatrix} 4 & 0 \\ 0 & 2 \end{bmatrix}

    The eigenvalues are λ1=4\lambda_1 = 4 and λ2=2\lambda_2 = 2. Both are positive, proving H≻0\mathbf{H} \succ 0 everywhere (strictly convex).

  3. First-order step with learning rate η=0.1\eta = 0.1:

    θ1=θ0−η∇L(θ0)=[00]−0.1[−4−2]=[0.40.2]\mathbf{\theta}_1 = \mathbf{\theta}_0 - \eta \nabla \mathcal{L}(\mathbf{\theta}_0) = \begin{bmatrix} 0 \\ 0 \end{bmatrix} - 0.1 \begin{bmatrix} -4 \\ -2 \end{bmatrix} = \begin{bmatrix} 0.4 \\ 0.2 \end{bmatrix}

    Initial loss: L(0,0)=5.0\mathcal{L}(0, 0) = 5.0. New loss:

    L(0.4,0.2)=2(0.4)2+(0.2)2−4(0.4)−2(0.2)+5=0.32+0.04−1.6−0.4+5=3.36\mathcal{L}(0.4, 0.2) = 2(0.4)^2 + (0.2)^2 - 4(0.4) - 2(0.2) + 5 = 0.32 + 0.04 - 1.6 - 0.4 + 5 = 3.36

    The loss drops from 5.05.0 to 3.363.36.

  4. Second-order Newton step: The inverse Hessian is:

    H−1=[4002]−1=[0.25000.5]\mathbf{H}^{-1} = \begin{bmatrix} 4 & 0 \\ 0 & 2 \end{bmatrix}^{-1} = \begin{bmatrix} 0.25 & 0 \\ 0 & 0.5 \end{bmatrix}

    The exact Newton displacement is:

    ΔθNewton=−H−1∇L(θ0)=−[0.25000.5][−4−2]=[1.01.0]\Delta \mathbf{\theta}_{\text{Newton}} = -\mathbf{H}^{-1} \nabla \mathcal{L}(\mathbf{\theta}_0) = -\begin{bmatrix} 0.25 & 0 \\ 0 & 0.5 \end{bmatrix} \begin{bmatrix} -4 \\ -2 \end{bmatrix} = \begin{bmatrix} 1.0 \\ 1.0 \end{bmatrix}

    New parameter: θNewton=[0,0]T+[1.0,1.0]T=[1.0,1.0]T\mathbf{\theta}_{\text{Newton}} = [0, 0]^T + [1.0, 1.0]^T = [1.0, 1.0]^T. Evaluating the gradient at this new point:

    ∇L(1.0,1.0)=[4(1.0)−42(1.0)−2]=[00]\nabla \mathcal{L}(1.0, 1.0) = \begin{bmatrix} 4(1.0) - 4 \\ 2(1.0) - 2 \end{bmatrix} = \begin{bmatrix} 0 \\ 0 \end{bmatrix}

    Newton's method reaches the exact global minimum in a single step with optimal loss L(1,1)=2(1)+1−4−2+5=2.0\mathcal{L}(1, 1) = 2(1) + 1 - 4 - 2 + 5 = 2.0.

Code

import numpy as np

def loss_and_grad(theta: np.ndarray) -> tuple[float, np.ndarray]:    """Compute quadratic loss and analytic gradient for theta = [theta_1, theta_2]."""    t1, t2 = theta[0], theta[1]    loss = float(2 * t1**2 + t2**2 - 4 * t1 - 2 * t2 + 5)    grad = np.array([4 * t1 - 4, 2 * t2 - 2], dtype=float)    return loss, grad

# Constant Hessian for quadratic losshessian = np.array([[4.0, 0.0], [0.0, 2.0]])inv_hessian = np.linalg.inv(hessian)
theta_0 = np.array([0.0, 0.0])l0, g0 = loss_and_grad(theta_0)print(f"Initial loss: {l0:.2f}, grad: {g0}")# -> Initial loss: 5.00, grad: [-4. -2.]
# 1. Gradient Descent Step (lr = 0.1)lr = 0.1theta_gd = theta_0 - lr * g0l_gd, g_gd = loss_and_grad(theta_gd)print(f"GD step: theta = {theta_gd}, loss = {l_gd:.2f}")# -> GD step: theta = [0.4 0.2], loss = 3.36
# 2. Newton's Method Steptheta_newton = theta_0 - inv_hessian @ g0l_newton, g_newton = loss_and_grad(theta_newton)print(f"Newton step: theta = {theta_newton}, loss = {l_newton:.2f}")# -> Newton step: theta = [1. 1.], loss = 2.00print(f"Newton residual gradient: {g_newton}")# -> Newton residual gradient: [0. 0.]

Watch Out For

Second-order Newton steps leaping toward saddle points and maxima

Newton's method seeks stationary points where ∇L=0\nabla \mathcal{L} = \mathbf{0}, but it is completely blind to whether that stationary point is a minimum or maximum. The update Δθ=−H−1∇L\Delta \mathbf{\theta} = -\mathbf{H}^{-1} \nabla \mathcal{L} inverts the Hessian unconditionally.

If the local Hessian possesses negative eigenvalues (λi<0\lambda_i < 0), the local curvature is convex downward. In those directions, Newton's update flips sign and steps directly toward the local maximum or saddle ridge, accelerating away from lower loss regions. Furthermore, computing and inverting an exact d×dd \times d Hessian matrix requires O(d3)O(d^3) arithmetic, which is intractable when d>105d > 10^5.

Fix: In high-dimensional non-convex optimization, prefer first-order methods with momentum (such as Adam), or use quasi-Newton trust-region methods (like L-BFGS with damping) that maintain positive-definite Hessian approximations and enforce descent conditions before taking a step.

The Quick Version

  • Calculus provides the gradient ∇L\nabla \mathcal{L} for local slope direction and the Hessian H\mathbf{H} for multivariable curvature.
  • Gradient descent follows the negative gradient direction with step size η\eta, trading second-order speed for linear O(d)O(d) computational efficiency per step.
  • Newton's method uses the inverse Hessian H−1\mathbf{H}^{-1} to eliminate curvature scale imbalances, reaching the minimum of a quadratic function in one step.
  • Stationary points where ∇L=0\nabla \mathcal{L} = \mathbf{0} are classified by Hessian eigenvalues: all positive indicates a minimum, all negative a maximum, and mixed signs a saddle point.
  • Modern deep learning relies predominantly on first-order stochastic methods because inverting large Hessians is computationally prohibitive and prone to saddle-point attraction.