Quantum Neural Networks
A quantum neural network arranges rotatable quantum gates into layered circuits analogous to deep neural networks, using classical gradient descent to steer quantum interference toward accurate predictions.
Why Does This Exist?
In classical deep learning, stacking linear layers with non-linear activations () yields deep hierarchical feature extractors capable of approximating arbitrary functions. However, classical models represent data as coordinate vectors in flat Euclidean space, which scales poorly when modeling highly correlated quantum systems, molecular spin configurations, or discrete combinatorial graphs.
Quantum Neural Networks (QNNs) bring the structural inductive bias of deep learning—depth, repeated modular layers, and localized parameter sharing—to parameterized quantum circuits. Instead of using arbitrary unstructured quantum circuits, a QNN stacks repeated blocks of single-qubit rotation gates (analogous to linear weights) and multi-qubit entangling gates (analogous to non-linear cross-channel interactions).
By embedding a QNN into a hybrid pipeline (for example, as a quantum layer between a classical convolutional front-end and a classical classification head), machine learning practitioners harness quantum expressive capacity. Stacking unitary layers generates Fourier-like frequency components that expand exponentially with circuit depth, enabling compact networks to model complex oscillatory or topological patterns with far fewer trainable parameters than classical MLPs.
Think of It Like This
A theatrical spotlight stage with motorized polarizing color filters
Imagine an auditorium where multiple beams of colored light travel across a stage toward a screen.
Along the path of the light beams are several consecutive racks of motorized optical filters. Each filter rack represents a layer of the neural network. Within each rack, servo motors rotate individual polarizing lenses by specific angles (trainable weights). Between racks, optical prisms cross and intertwine adjacent light beams, causing them to physically overlap and interfere (entanglement).
At the back wall, light sensors measure the final color intensities and brightness at discrete spots (quantum measurement readout). A lighting director sitting in the control booth (the classical optimizer) checks whether the final stage pattern matches the desired dramatic mood. The director computes how much each filter contributed to the error, sends servo signals back to rotate each individual filter angle slightly, and repeats the test until the interference pattern is dialed in.
How It Actually Works
Layered Ansatz Construction and Hybrid Backpropagation
A Quantum Neural Network processes data through an alternating sequence of parameterized single-qubit transformations and multi-qubit entangling operations organized into discrete layers:
Each layer is partitioned into two fundamental sub-blocks:
- Parameterized Single-Qubit Rotations: Each qubit undergoes an arbitrary 3D rotation on the Bloch sphere parameterized by Euler angles :
- Entangling Fabric: Fixed two-qubit gates (such as CNOT or CZ) entangle adjacent qubits in a 1D chain or 2D grid. Entanglement spreads information globally across the -dimensional Hilbert space:
At the output of layer , the network extracts a vector of classical continuous features by measuring the expectation values of local Pauli observables :
These quantum expectations can then be piped directly into a classical linear classification head:
Hybrid Backpropagation via the Chain Rule
When optimizing the total cross-entropy loss , gradients with respect to classical parameters () are calculated via standard automatic differentiation. To compute gradients with respect to the internal quantum gate parameters , the system uses the multivariable chain rule:
The classical derivative is supplied by PyTorch's autograd engine, while the quantum sensitivity is evaluated on the quantum processor using the parameter-shift rule:
Worked Example
Consider a 2-qubit QNN with a single parameterized layer.
- Initial state:
- Input encoding: Qubit 0 receives , Qubit 1 receives via :
Step 1: Entangling Operation (CNOT) Apply CNOT with Qubit 0 as control and Qubit 1 as target: Because Qubit 0 is in both basis terms, the target Qubit 1 is unaffected:
Step 2: Parameterized Rotation on Qubit 1 Apply on Qubit 1 with current parameter : Notice that Qubit 1 is currently at angle . Adding another rotation of rotates Qubit 1 to a total angle of :
The state of the 2-qubit system becomes purely:
Step 3: Measure Observable on Qubit 1 The Pauli-Z matrix acts as and . Because Qubit 1 is in state , the expectation value is:
Step 4: Compute Loss and Parameter Gradient Suppose our training target is under mean squared error :
Classical gradient factor: . Now evaluate quantum sensitivity using parameter shifts at :
- Shift up: . .
- Shift down: . .
- Parameter shift gradient:
Total chain-rule gradient: . At , the expectation reaches its global minimum (), correctly registering zero gradient slope.
Code
The following Python script simulates a 2-qubit Quantum Neural Network layer coupled with a classical loss function:
import numpy as np
def ry(theta: float) -> np.ndarray: half = theta / 2.0 return np.array([ [np.cos(half), -np.sin(half)], [np.sin(half), np.cos(half)] ], dtype=np.complex128)
cnot = np.array([ [1, 0, 0, 0], [0, 1, 0, 0], [0, 0, 0, 1], [0, 0, 1, 0]], dtype=np.complex128)
z_obs = np.array([[1, 0], [0, -1]], dtype=np.complex128)identity_2 = np.eye(2, dtype=np.complex128)z1_observable = np.kron(identity_2, z_obs) # I (x) Z
def forward_qnn(x_angles: list[float], theta: float) -> float: """Simulates 2-qubit QNN forward pass: returns expectation <Z_1>.""" # State preparation |00> state = np.array([1, 0, 0, 0], dtype=np.complex128)
# Input encoding: Ry(x0) (x) Ry(x1) enc = np.kron(ry(x_angles[0]), ry(x_angles[1])) state = enc @ state
# Entangling layer: CNOT state = cnot @ state
# Parameterized rotation on Qubit 1: I (x) Ry(theta) u_param = np.kron(identity_2, ry(theta)) state = u_param @ state
# Measurement exp_val = np.real(np.conj(state).T @ z1_observable @ state) return float(exp_val)
def hybrid_backward(x_angles: list[float], theta: float, target: float) -> tuple[float, float]: """Computes hybrid loss and exact chain-rule gradient dL/dtheta.""" y = forward_qnn(x_angles, theta) loss = 0.5 * (y - target) ** 2 dl_dy = y - target
# Parameter-shift rule for dy/dtheta shift = np.pi / 2.0 y_plus = forward_qnn(x_angles, theta + shift) y_minus = forward_qnn(x_angles, theta - shift) dy_dtheta = 0.5 * (y_plus - y_minus)
dl_dtheta = dl_dy * dy_dtheta return loss, dl_dtheta
# Test parameters matching worked examplex_inputs = [0.0, np.pi / 2.0]theta_weight = np.pi / 2.0target_y = 1.0
loss_val, grad_val = hybrid_backward(x_inputs, theta_weight, target_y)
print(f"Readout y: {forward_qnn(x_inputs, theta_weight):.4f}")# -> Readout y: -1.0000print(f"Loss: {loss_val:.4f}")# -> Loss: 2.0000print(f"Hybrid gradient dL/dtheta: {grad_val:.4f}")# -> Hybrid gradient dL/dtheta: 0.0000Watch Out For
Shot noise variance corrupting classical gradient descent
On physical quantum processors, expectation values cannot be computed analytically via statevector products. Instead, the quantum computer executes the circuit repeatedly for runs (typically 1,000 to 10,000 shots) and computes the sample mean of binary detector clicks.
Because quantum measurement is fundamentally probabilistic, any expectation value estimate carries statistical shot noise with variance . When computing the parameter-shift gradient , subtracting two noisy estimates doubles the variance. Standard classical optimizers like Adam or SGD will misinterpret this stochastic shot noise as genuine loss curvature, causing parameter updates to oscillate violently around local optima.
Fix: Never use standard vanilla SGD with fixed learning rates on physical QPUs. Employ Simultaneous Perturbation Stochastic Approximation (SPSA), which estimates gradients using random perturbations robust to shot noise, or implement adaptive shot budgeting, where the number of measurement shots dynamically increases as the optimizer approaches convergence.
The Quick Version
- Quantum Neural Networks organize parameterized quantum gates into deep, layered architectures composed of rotation and entanglement stages.
- The hybrid quantum-classical pipeline combines quantum circuits with classical neural networks via joint backpropagation.
- The multivariable chain rule combines PyTorch classical autograd with the quantum parameter-shift rule.
- Physical execution relies on finite measurement shots, requiring noise-aware optimizers like SPSA to withstand statistical measurement variance.