Skip to content
AI360Xpert
Beta
Math, Statistics & Probability for ML

Math, Statistics & Probability for ML

30 interview questions in this topic, each with its full answer shown below. Use "Collapse all" to skim just the titles.

“How would you design an A/B test with low power or a heavy-tailed metric, and what statistical tools would you use?”

Quick answer

For heavy-tailed metrics, clip outliers, apply log transformations, or use non-parametric tests like Mann-Whitney. To fix low power, increase sample size or use variance reduction techniques like CUPED.

💡 Note For heavy-tailed metrics, clip outliers, apply log transformations, or use non-parametric tests like Mann-Whitney. To fix low power, increase sample size or use variance reduction techniques like CUPED.

Answer

Designing A/B tests in scenarios with low statistical power or highly skewed, heavy-tailed metrics (like revenue per user or time on page) requires specialized techniques, as standard t-tests will fail or require impractically large sample sizes.

Handling Heavy-Tailed Metrics:

  • Outlier Capping: Clip extreme values at a specific percentile (e.g., 99th) to prevent single massive observations from skewing the mean.
  • Transformations: Apply a log-transformation log⁡(x+1)\log(x+1) to compress the long tail into a more normal distribution.
  • Non-Parametric Tests: Use the Mann-Whitney U test, which compares ranks rather than means, making it highly robust to outliers.

Increasing Power:

  • Variance Reduction: Use techniques like CUPED (Controlled-experiment Using Pre-Experiment Data). CUPED uses pre-experiment historical data as a covariate to adjust the target metric, stripping away inherent user variance and dramatically shrinking confidence intervals.
  • Proxy Metrics: If the true goal (e.g., a purchase) is rare, test against a higher-frequency leading proxy metric (e.g., adding to cart) to increase the sample size of positive events.

“What is Bayes' theorem?”

Quick answer

Bayes' theorem updates the probability of a hypothesis based on new evidence. It relates conditional probabilities as P(A|B) = P(B|A)*P(A) / P(B).

💡 Note Bayes' theorem updates the probability of a hypothesis based on new evidence. It relates conditional probabilities as P(A|B) = P(B|A)*P(A) / P(B).

Answer

Bayes' theorem is a fundamental principle in probability that describes how to update the probabilities of hypotheses when given evidence. It links the degree of belief in a proposition before and after accounting for evidence.

The formula is: P(A∣B)=P(B∣A)P(A)P(B)P(A|B) = \frac{P(B|A)P(A)}{P(B)} Where:

  • P(A)P(A) is the prior: initial belief about AA.
  • P(B∣A)P(B|A) is the likelihood: probability of observing evidence BB if AA is true.
  • P(B)P(B) is the marginal likelihood (evidence).
  • P(A∣B)P(A|B) is the posterior: updated belief about AA given BB.

In ML, Bayes' theorem is the foundation of Bayesian inference, Naive Bayes classifiers, and Maximum A Posteriori (MAP) estimation, allowing models to incorporate prior knowledge.

“How does the bias of an estimator relate to its variance and mean squared error?”

Quick answer

The Mean Squared Error of an estimator can be decomposed into the sum of its squared Bias and its Variance. This illustrates the fundamental Bias-Variance Tradeoff.

💡 Note The Mean Squared Error of an estimator can be decomposed into the sum of its squared Bias and its Variance. This illustrates the fundamental Bias-Variance Tradeoff.

Answer

In statistics and machine learning, we often use an estimator θ^\hat{\theta} to approximate a true parameter θ\theta. The quality of this estimator is usually measured by its Mean Squared Error (MSE), defined as E[(θ^−θ)2]E[(\hat{\theta} - \theta)^2].

The MSE can be mathematically decomposed into two distinct components: MSE = Bias(θ^)2+Variance(θ^)Bias(\hat{\theta})^2 + Variance(\hat{\theta})

  • Bias is the error introduced by approximating a real-world problem, which may be complex, by a simplified model. It is the difference between the expected value of our estimator and the true value: E[θ^]−θE[\hat{\theta}] - \theta. High bias implies underfitting.
  • Variance is the amount the estimator θ^\hat{\theta} changes if estimated using different training datasets. High variance implies the model is highly sensitive to statistical noise in the training data (overfitting).

This decomposition forms the basis of the Bias-Variance Tradeoff. As you increase model complexity, bias decreases but variance increases. The goal of hyperparameter tuning (like adjusting regularization) is to find the sweet spot that minimizes the overall MSE.

“What is the Central Limit Theorem, and why does it matter for ML?”

Quick answer

The CLT states that the sum or average of many independent random variables tends toward a normal distribution, regardless of the original distribution. This justifies assuming Gaussian errors and simplifies ML models.

💡 Note The CLT states that the sum or average of many independent random variables tends toward a normal distribution, regardless of the original distribution. This justifies assuming Gaussian errors and simplifies ML models.

Answer

The Central Limit Theorem (CLT) states that when independent random variables are added together, their properly normalized sum tends toward a normal (Gaussian) distribution, even if the original variables themselves are not normally distributed, provided the sample size is sufficiently large.

This theorem is deeply important in machine learning and statistics for several reasons:

  • Assumption Justification: Because errors or noise in real-world data are often the sum of many independent, unobserved random effects, the CLT justifies modeling these errors using a Gaussian distribution (e.g., in linear regression).
  • Hypothesis Testing: It allows us to use normal distribution properties for sample means in A/B testing and confidence intervals, even when the underlying population distribution is unknown or heavily skewed.
  • Stochastic Optimization: The noise in stochastic gradient descent (SGD) often behaves roughly Gaussian over large mini-batches due to the CLT.

“What is the chain rule, and how does it underpin backpropagation?”

Quick answer

The chain rule is a calculus formula for computing the derivative of composed functions. Backpropagation applies the chain rule repeatedly to compute gradients in neural networks.

💡 Note The chain rule is a calculus formula for computing the derivative of composed functions. Backpropagation applies the chain rule repeatedly to compute gradients in neural networks.

Answer

The chain rule is a fundamental theorem in calculus used to compute the derivative of a composite function. If you have a function f(g(x))f(g(x)), the chain rule states that its derivative with respect to xx is f′(g(x))⋅g′(x)f'(g(x)) \cdot g'(x). In Leibniz notation for z=f(y)z = f(y) and y=g(x)y = g(x), it is written as dzdx=dzdy⋅dydx\frac{dz}{dx} = \frac{dz}{dy} \cdot \frac{dy}{dx}.

Backpropagation is the algorithm used to train neural networks, and it is essentially a highly efficient, organized application of the chain rule. A neural network is a massive composite function: L(W3σ(W2σ(W1x)))L(W_3 \sigma(W_2 \sigma(W_1 x))).

To update the weights W1W_1 in an early layer, backpropagation calculates the gradient of the loss LL with respect to W1W_1 by multiplying the local gradients of each layer, working backward from the output layer to the input. Without the chain rule, computing these gradients analytically for millions of parameters would be computationally intractable.

“What is conditional probability?”

Quick answer

Conditional probability is the probability of an event occurring given that another event has already occurred, calculated as P(A|B) = P(A and B) / P(B).

💡 Note Conditional probability is the probability of an event occurring given that another event has already occurred, calculated as P(A|B) = P(A and B) / P(B).

Answer

Conditional probability measures the likelihood of an event AA happening, assuming that another event BB has already occurred. It restricts the sample space to the outcomes where BB happens.

The formula is P(A∣B)=P(A∩B)P(B)P(A|B) = \frac{P(A \cap B)}{P(B)}, where P(A∩B)P(A \cap B) is the joint probability of both events happening, and P(B)>0P(B) > 0.

In machine learning, conditional probability is everywhere. For instance, in classification, we want to find P(y∣X)P(y|X) — the conditional probability of class yy given the input features XX. Generative models, Markov chains, and Bayesian networks all fundamentally rely on modeling conditional probabilities.

“How do confidence intervals differ from credible intervals?”

Quick answer

A confidence interval is a frequentist concept where the bounds are random variables that contain the fixed true parameter in X% of samples. A credible interval is a Bayesian concept where the true parameter has an X% probability of being within the fixed bounds.

💡 Note A confidence interval is a frequentist concept where the bounds are random variables that contain the fixed true parameter in X% of samples. A credible interval is a Bayesian concept where the true parameter has an X% probability of being within the fixed bounds.

Answer

The difference between these intervals highlights the core philosophical split between Frequentist and Bayesian statistics.

A Confidence Interval (Frequentist) assumes the true parameter is a fixed, unknown value. The interval itself is considered random (it changes with each sample). A 95% confidence interval means that if you repeated the experiment infinitely many times and generated intervals for each, 95% of those calculated intervals would contain the true parameter. You cannot say there is a 95% chance the parameter is in this specific interval.

A Credible Interval (Bayesian) treats the true parameter as a random variable with a probability distribution. Based on the posterior distribution (after observing data and applying priors), a 95% credible interval simply means there is a 95% probability that the true parameter lies within that specific interval.

Credible intervals are generally much more intuitive to non-statisticians, as they directly answer "what is the probability the parameter is in this range?"

“What is the difference between convex and non-convex optimization?”

Quick answer

Convex optimization deals with functions having a single global minimum, guaranteeing convergence. Non-convex optimization deals with landscapes containing multiple local minima and saddle points, common in deep learning.

💡 Note Convex optimization deals with functions having a single global minimum, guaranteeing convergence. Non-convex optimization deals with landscapes containing multiple local minima and saddle points, common in deep learning.

Answer

A function is convex if a line segment drawn between any two points on its graph lies above or on the graph. In optimization, this creates a bowl-like landscape. The defining feature of convex optimization is that any local minimum is guaranteed to be the global minimum. Problems like Linear Regression, Logistic Regression, and Support Vector Machines (without certain kernels) are convex. Gradient descent on these functions is highly reliable.

Non-convex functions have wavy, irregular landscapes filled with multiple local minima, plateaus, and saddle points. Deep neural networks inherently produce highly non-convex loss landscapes due to nonlinear activations and layered architectures.

In non-convex optimization, standard gradient descent can easily get stuck in local minima or slow down at saddle points. This necessitates advanced techniques like momentum, adaptive learning rates (Adam), and stochasticity (SGD) to help the optimizer escape suboptimal regions and find a "good enough" minimum.

“What is correlation, and why doesn't it imply causation?”

Quick answer

Correlation measures the statistical relationship or dependence between two variables. It doesn't imply causation because an observed relationship could be due to a hidden confounding variable or sheer coincidence.

💡 Note Correlation measures the statistical relationship or dependence between two variables. It doesn't imply causation because an observed relationship could be due to a hidden confounding variable or sheer coincidence.

Answer

Correlation is a statistical measure that expresses the extent to which two variables are linearly related (meaning they change together at a constant rate). The most common metric is Pearson's correlation coefficient, ranging from -1 to 1.

However, correlation does not imply causation. Just because two variables move together doesn't mean one causes the other. This can happen due to:

  1. Confounding Variables: A third, unseen variable might be causing both. For example, ice cream sales and shark attacks are highly correlated, but both are caused by a confounding variable: summer heat.
  2. Reverse Causation: BB might cause AA rather than AA causing BB.
  3. Spurious Correlation: In small samples or massive datasets, variables can appear highly correlated purely by random chance.

To prove causation, one typically needs controlled, randomized experiments or advanced causal inference techniques.

“What is the difference between covariance and correlation?”

Quick answer

Covariance measures the direction of the linear relationship between two variables, but is scale-dependent. Correlation is normalized covariance, providing both direction and a standardized strength (from -1 to 1).

💡 Note Covariance measures the direction of the linear relationship between two variables, but is scale-dependent. Correlation is normalized covariance, providing both direction and a standardized strength (from -1 to 1).

Answer

Covariance indicates the direction of the linear relationship between two random variables. If both variables tend to increase together, covariance is positive; if one increases while the other decreases, it is negative. However, the exact value of covariance is hard to interpret because it depends heavily on the units (scale) of the variables.

Correlation (specifically Pearson correlation) is a scaled version of covariance. By dividing the covariance by the product of the standard deviations of the two variables, correlation becomes dimensionless. It always falls between -1 and 1.

While covariance is useful for matrix operations like in the Multivariate Normal Distribution or PCA, correlation is much more practical for quickly interpreting the strength and direction of a relationship between two features in exploratory data analysis.

“Derive the gradient of the softmax cross-entropy loss.”

Quick answer

The gradient of the softmax cross-entropy loss with respect to the pre-activation logits is simply (p - y), where p is the predicted softmax probability and y is the one-hot target label.

💡 Note The gradient of the softmax cross-entropy loss with respect to the pre-activation logits is simply (p - y), where p is the predicted softmax probability and y is the one-hot target label.

Answer

When combining the Softmax activation function with categorical Cross-Entropy loss, the resulting gradient with respect to the input logits is remarkably elegant.

Let the true one-hot label vector be yy and the input logits be zz. The predicted probabilities are pi=ezi∑jezjp_i = \frac{e^{z_i}}{\sum_j e^{z_j}}. The cross-entropy loss is L=−∑kyklog⁡(pk)L = -\sum_k y_k \log(p_k).

To find ∂L∂zi\frac{\partial L}{\partial z_i}, we apply the chain rule. The derivative of the softmax output pkp_k with respect to ziz_i is pi(1−pi)p_i(1 - p_i) if k=ik=i, and −pkpi-p_k p_i if k≠ik \neq i.

Plugging this into the loss derivative: ∂L∂zi=−∑kyk1pk∂pk∂zi\frac{\partial L}{\partial z_i} = -\sum_k y_k \frac{1}{p_k} \frac{\partial p_k}{\partial z_i} =−yi1pipi(1−pi)−∑k≠iyk1pk(−pkpi)= - y_i \frac{1}{p_i} p_i(1 - p_i) - \sum_{k \neq i} y_k \frac{1}{p_k} (-p_k p_i) =−yi+yipi+∑k≠iykpi= -y_i + y_i p_i + \sum_{k \neq i} y_k p_i =−yi+pi∑kyk= -y_i + p_i \sum_k y_k

Since yy is one-hot, ∑kyk=1\sum_k y_k = 1. The expression simplifies beautifully to: ∂L∂zi=pi−yi\frac{\partial L}{\partial z_i} = p_i - y_i

This simple difference between predicted probability and true label makes the backward pass incredibly efficient.

“What are eigenvalues and eigenvectors, and how do they relate to PCA?”

Quick answer

Eigenvectors are vectors whose direction is unchanged by a linear transformation, and eigenvalues are their scaling factors. In PCA, the eigenvectors of the covariance matrix form the principal components.

💡 Note Eigenvectors are vectors whose direction is unchanged by a linear transformation, and eigenvalues are their scaling factors. In PCA, the eigenvectors of the covariance matrix form the principal components.

Answer

An eigenvector of a square matrix AA is a non-zero vector v\mathbf{v} that only changes by a scalar factor when the linear transformation AA is applied to it. This scalar factor is the corresponding eigenvalue λ\lambda. The relationship is defined as Av=λvA\mathbf{v} = \lambda\mathbf{v}.

In Principal Component Analysis (PCA), we compute the covariance matrix of the data, which captures how features vary together. By finding the eigenvectors and eigenvalues of this covariance matrix, we discover the underlying structure of the data.

The eigenvectors represent the directions of maximum variance (the principal components), and their corresponding eigenvalues indicate the magnitude of the variance in those directions. By selecting the top eigenvectors with the largest eigenvalues, PCA projects the data into a lower-dimensional space while retaining the most critical information.

“Explain the Expectation-Maximization algorithm and why it converges.”

Quick answer

EM is an iterative algorithm to find maximum likelihood estimates in models with latent variables. It alternates between estimating the hidden variables (E-step) and optimizing the parameters (M-step).

💡 Note EM is an iterative algorithm to find maximum likelihood estimates in models with latent variables. It alternates between estimating the hidden variables (E-step) and optimizing the parameters (M-step).

Answer

The Expectation-Maximization (EM) algorithm is a powerful iterative method used for finding Maximum Likelihood Estimates (MLE) of parameters in statistical models where the data is incomplete or has unobserved (latent) variables. A classic application is the Gaussian Mixture Model (GMM).

The algorithm repeatedly alternates between two steps:

  1. E-Step (Expectation): Calculate the expected value of the latent variables using the current estimate of the parameters. In GMMs, this means calculating the probability that each point belongs to each cluster.
  2. M-Step (Maximization): Update the model parameters to maximize the expected log-likelihood found in the E-step, assuming the latent variables are exactly what we estimated.

Convergence: EM is guaranteed to converge to a local maximum (or saddle point) because the likelihood is monotonically non-decreasing after every single iteration. By bounding the log-likelihood function from below using Jensen's inequality and maximizing that lower bound, the overall likelihood never drops.

“Explain the role of the Fisher information matrix and natural gradients.”

Quick answer

The Fisher information matrix measures how much information an observable carries about a parameter. Natural gradients use the Fisher inverse to correct gradients, moving them along the probabilistic manifold rather than flat Euclidean space.

💡 Note The Fisher information matrix measures how much information an observable carries about a parameter. Natural gradients use the Fisher inverse to correct gradients, moving them along the probabilistic manifold rather than flat Euclidean space.

Answer

The Fisher Information Matrix (FIM) is a concept from information theory and statistics that quantifies the amount of information that an observable random variable carries about an unknown parameter. In maximum likelihood estimation, the FIM describes the local curvature of the log-likelihood function. Interestingly, under certain conditions, it is equivalent to the Hessian of the KL divergence.

In standard neural network optimization, standard Gradient Descent operates in Euclidean space. It doesn't respect the fact that deep learning models output probability distributions (like softmax). A small step in parameter space might drastically alter the output distribution.

Natural Gradient Descent solves this by pre-multiplying the standard gradient by the inverse of the Fisher Information Matrix: Δθ=−F−1∇L(θ)\Delta \theta = -F^{-1} \nabla L(\theta). This reshapes the step size to move along the Riemann manifold of probability distributions. By maintaining a constant KL divergence step rather than a constant Euclidean distance step, natural gradients provide much faster, curvature-aware convergence, avoiding pathological plateaus.

“What is a gradient?”

Quick answer

A gradient is a vector containing all the partial derivatives of a multivariable function. It points in the direction of the steepest ascent of the function.

💡 Note A gradient is a vector containing all the partial derivatives of a multivariable function. It points in the direction of the steepest ascent of the function.

Answer

A gradient is a generalization of the concept of a derivative to functions of several variables. If you have a scalar-valued multivariable function f(x1,x2,...,xn)f(x_1, x_2, ..., x_n), its gradient, denoted as ∇f\nabla f, is a vector consisting of all its partial derivatives.

Geometrically, the gradient vector at any given point points in the direction of the steepest ascent of the function, and its magnitude represents the rate of increase in that direction.

In machine learning, the gradient is the core of optimization algorithms like Gradient Descent. To minimize a loss function L(θ)L(\theta) with respect to the model parameters θ\theta, we compute the gradient ∇L(θ)\nabla L(\theta) and update the parameters by taking a step in the opposite direction of the gradient (steepest descent).

“Why does gradient descent converge on convex functions, and what rate can you expect?”

Quick answer

Gradient descent converges on convex functions because they have a single global minimum and gradients constantly point toward it. With a suitable learning rate, convergence rates are typically O(1/t).

💡 Note Gradient descent converges on convex functions because they have a single global minimum and gradients constantly point toward it. With a suitable learning rate, convergence rates are typically O(1/t).

Answer

For convex functions, the loss landscape resembles a smooth bowl. The defining mathematical property of convexity ensures that the gradient always points generally in the direction of the single global minimum. There are no local minima or saddle points to trap the optimizer.

Gradient descent updates parameters via θt+1=θt−η∇f(θt)\theta_{t+1} = \theta_t - \eta \nabla f(\theta_t). If the function is sufficiently smooth (specifically, its gradient is Lipschitz continuous), and the learning rate η\eta is small enough, each step is guaranteed to decrease the loss monotonically.

Convergence Rates:

  • For standard convex functions, gradient descent achieves an error rate of O(1/t)O(1/t) after tt iterations.
  • If the function is strongly convex (it has a guaranteed minimum curvature), the convergence becomes exponentially faster, achieving a rate of O(e−αt)O(e^{-\alpha t}), also known as linear convergence. While it guarantees convergence, choosing an appropriate learning rate is crucial—too high causes divergence, while too low causes painfully slow progress.

“What is the Jacobian and the Hessian, and where do they appear in ML?”

Quick answer

The Jacobian is the matrix of first-order partial derivatives of a vector-valued function. The Hessian is the square matrix of second-order partial derivatives of a scalar function.

💡 Note The Jacobian is the matrix of first-order partial derivatives of a vector-valued function. The Hessian is the square matrix of second-order partial derivatives of a scalar function.

Answer

The Jacobian matrix contains all first-order partial derivatives of a vector-valued function. If a function maps Rn\mathbb{R}^n to Rm\mathbb{R}^m, its Jacobian is an m×nm \times n matrix. In machine learning, Jacobians are heavily utilized during backpropagation to compute the gradients of layer outputs with respect to layer inputs or weights, especially when the activation function outputs a vector (like softmax).

The Hessian matrix contains all second-order partial derivatives of a scalar-valued function. For a function mapping Rn\mathbb{R}^n to R\mathbb{R}, the Hessian is an n×nn \times n symmetric matrix. The Hessian describes the local curvature of the loss function landscape.

In optimization, while first-order methods (like Gradient Descent) only use the gradient, second-order methods (like Newton's method) use the inverse of the Hessian to take much more informed steps, adapting to the curvature. However, computing and storing the Hessian is often too expensive for large neural networks, leading to approximations like L-BFGS.

“What is KL divergence, and how does it differ from cross-entropy?”

Quick answer

KL divergence measures the difference between two probability distributions. Cross-entropy is the sum of the entropy of the true distribution and the KL divergence, often used as a loss function.

💡 Note KL divergence measures the difference between two probability distributions. Cross-entropy is the sum of the entropy of the true distribution and the KL divergence, often used as a loss function.

Answer

Kullback-Leibler (KL) Divergence is an information-theoretic measure of how one probability distribution QQ differs from a reference distribution PP. It measures the expected excess surprise or "information loss" when using QQ to approximate PP. It is not symmetric (i.e., DKL(P∣∣Q)≠DKL(Q∣∣P)D_{KL}(P||Q) \neq D_{KL}(Q||P)).

Cross-Entropy, H(P,Q)H(P, Q), measures the total average number of bits needed to encode events from PP using an optimized code designed for QQ. The relationship is: H(P,Q)=H(P)+DKL(P∣∣Q)H(P, Q) = H(P) + D_{KL}(P||Q), where H(P)H(P) is the intrinsic entropy of PP.

In machine learning classification, PP is the fixed true label distribution (often one-hot encoded, making H(P)=0H(P)=0), and QQ is the model's predicted distribution. Because H(P)H(P) is constant, minimizing Cross-Entropy is mathematically equivalent to minimizing KL divergence, bringing QQ as close to PP as possible.

“What is a matrix multiplication, and what does it represent?”

Quick answer

Matrix multiplication is an operation that takes two matrices and produces a third. It represents the composition of linear transformations.

💡 Note Matrix multiplication is an operation that takes two matrices and produces a third. It represents the composition of linear transformations.

Answer

Matrix multiplication is a binary operation that produces a matrix from two matrices. To multiply an m×nm \times n matrix AA by an n×pn \times p matrix BB, you compute the dot products of the rows of AA with the columns of BB, resulting in an m×pm \times p matrix CC.

Conceptually, every matrix represents a linear transformation (like rotation, scaling, or shearing). Multiplying two matrices is mathematically equivalent to composing their respective linear transformations—applying one transformation and then the other.

In Deep Learning, matrix multiplication is the backbone of forward and backward passes in neural networks. The transformation of activations from one layer to the next via weight matrices is exactly a large-scale matrix multiplication, which is why GPUs (optimized for this operation) are essential for ML.

“How do MCMC methods like Metropolis-Hastings sample from complex distributions?”

Quick answer

MCMC constructs a Markov chain that wanders through the state space. Metropolis-Hastings generates proposals and accepts them probabilistically, ensuring the stationary distribution matches the target complex distribution.

💡 Note MCMC constructs a Markov chain that wanders through the state space. Metropolis-Hastings generates proposals and accepts them probabilistically, ensuring the stationary distribution matches the target complex distribution.

Answer

Markov Chain Monte Carlo (MCMC) methods are a class of algorithms used to sample from probability distributions that are too complex to evaluate directly, particularly the unnormalized posterior distributions in Bayesian inference.

The core idea is to design a Markov Chain (a random walk where the next state depends only on the current state) whose stationary distribution exactly matches the target distribution we want to sample from.

The Metropolis-Hastings (MH) algorithm is the most famous MCMC method. It works as follows:

  1. Start at a random initial state xx.
  2. Propose a new state x′x' using a simple proposal distribution (e.g., adding Gaussian noise to xx).
  3. Calculate an acceptance ratio α\alpha, which compares the probability of the proposed state against the current state, adjusted for the proposal distribution's asymmetry.
  4. Accept the new state x′x' with probability min⁡(1,α)\min(1, \alpha). If rejected, stay at xx.

Over many iterations, the chain spends more time in regions of high probability, yielding samples that accurately represent the target distribution.

“What are mean, median, and mode, and when does each mislead?”

Quick answer

Mean is the average, median is the middle value, and mode is the most frequent. Mean misleads with outliers, median ignores exact values, and mode can be uninformative for continuous data.

💡 Note Mean is the average, median is the middle value, and mode is the most frequent. Mean misleads with outliers, median ignores exact values, and mode can be uninformative for continuous data.

Answer

The mean (average) is calculated by summing all values and dividing by the count. It is sensitive to outliers and can heavily skew if extreme values are present. The median is the middle value when the data is sorted. It is robust to outliers but ignores the magnitude of most values, which can hide important changes in the data's tails. The mode is the most frequent value. While useful for categorical data, it can be meaningless for continuous distributions where every value is unique.

In machine learning, choosing the right measure of central tendency is critical for data imputation and scaling. Replacing missing values with the mean in a skewed feature might introduce bias, whereas the median would provide a more representative typical value.

“Explain maximum likelihood estimation versus maximum a posteriori estimation.”

Quick answer

MLE finds parameters that maximize the likelihood of the observed data. MAP does the same but includes a prior distribution over the parameters, acting as a regularizer.

💡 Note MLE finds parameters that maximize the likelihood of the observed data. MAP does the same but includes a prior distribution over the parameters, acting as a regularizer.

Answer

Maximum Likelihood Estimation (MLE) is a method to estimate the parameters of a statistical model. It finds the parameter values θ\theta that maximize the likelihood function P(Data∣θ)P(Data | \theta) — meaning it chooses parameters that make the observed data most probable.

Maximum A Posteriori (MAP) estimation extends MLE by incorporating Bayesian prior beliefs. Instead of maximizing the likelihood, MAP maximizes the posterior distribution P(θ∣Data)∝P(Data∣θ)P(θ)P(\theta | Data) \propto P(Data | \theta) P(\theta). Here, P(θ)P(\theta) is the prior distribution over the parameters.

In machine learning, MLE is equivalent to minimizing the standard loss functions (like Mean Squared Error for Gaussian noise, or Cross-Entropy for categorical data). MAP adds a regularization term. For instance, assuming a Gaussian prior on the weights in linear regression yields L2 regularization (Ridge regression), while a Laplace prior yields L1 regularization (Lasso).

“Explain the multiple testing problem and corrections such as Bonferroni and FDR.”

Quick answer

The multiple testing problem occurs when testing many hypotheses simultaneously, drastically increasing the chance of false positives. Corrections like Bonferroni adjust the significance threshold to control this error.

💡 Note The multiple testing problem occurs when testing many hypotheses simultaneously, drastically increasing the chance of false positives. Corrections like Bonferroni adjust the significance threshold to control this error.

Answer

The Multiple Testing Problem arises when you perform many statistical hypothesis tests simultaneously. If you test a single hypothesis at a significance level of α=0.05\alpha = 0.05, there is a 5% chance of a false positive (Type I error). However, if you test 100 features simultaneously, the probability of getting at least one false positive skyrockets to 1−(1−0.05)100≈99.4%1 - (1 - 0.05)^{100} \approx 99.4\%.

To prevent spurious discoveries (common in A/B testing or genomics), we must apply corrections:

  1. Bonferroni Correction: The strictest approach. It controls the Family-Wise Error Rate (FWER) by dividing the significance level by the number of tests (mm): αnew=α/m\alpha_{new} = \alpha / m. While it guarantees a low false positive rate, it severely reduces statistical power, causing many false negatives.
  2. False Discovery Rate (FDR) / Benjamini-Hochberg: A more modern, less conservative approach. Instead of preventing any false positives, it controls the expected proportion of false discoveries among the rejected hypotheses. This maintains much higher power, making it preferable for large-scale machine learning and data mining.

“What is mutual information, and how is it used in feature selection?”

Quick answer

Mutual information measures the amount of information obtained about one variable through observing another. It is used in feature selection because it captures non-linear dependencies, unlike correlation.

💡 Note Mutual information measures the amount of information obtained about one variable through observing another. It is used in feature selection because it captures non-linear dependencies, unlike correlation.

Answer

Mutual Information (MI) is a concept from information theory that quantifies the "amount of information" (in bits or shannons) one random variable contains about another. Mathematically, it is the KL divergence between the joint distribution and the product of the marginal distributions: I(X;Y)=DKL(P(X,Y)∣∣P(X)P(Y))I(X; Y) = D_{KL}(P(X,Y) || P(X)P(Y)). If XX and YY are strictly independent, MI is zero.

In machine learning, MI is a highly effective metric for feature selection (specifically, filter methods).

Unlike Pearson correlation, which only measures linear relationships, Mutual Information can capture purely non-linear dependencies. By calculating the MI between each input feature and the target label, you can rank the features based on how much predictive information they carry. Selecting the top-K features with the highest MI helps reduce dimensionality while retaining the most discriminative signals for the model.

“What is a p-value?”

Quick answer

A p-value is the probability of observing results at least as extreme as the observed ones, assuming the null hypothesis is true.

💡 Note A p-value is the probability of observing results at least as extreme as the observed ones, assuming the null hypothesis is true.

Answer

A p-value is a foundational concept in frequentist hypothesis testing. It quantifies the probability of obtaining test results at least as extreme as the ones observed during the experiment, assuming that the null hypothesis (H0H_0) is correct.

A small p-value (typically ≤0.05\le 0.05) suggests that the observed data is very unlikely under the null hypothesis, providing evidence to reject H0H_0 in favor of an alternative hypothesis. Conversely, a large p-value implies that the data is consistent with the null hypothesis.

It is a common misconception that the p-value is the probability that the null hypothesis is true. It is not. It strictly measures the compatibility of the data with the null hypothesis. In machine learning, p-values are often used in A/B testing, feature selection, and evaluating the statistical significance of model coefficients.

“What is a probability distribution? Give examples of common ones.”

Quick answer

A probability distribution describes how probabilities are distributed over values of a random variable. Common examples include Normal, Binomial, and Poisson distributions.

💡 Note A probability distribution describes how probabilities are distributed over values of a random variable. Common examples include Normal, Binomial, and Poisson distributions.

Answer

A probability distribution is a mathematical function that provides the probabilities of occurrence of different possible outcomes in an experiment. For a discrete random variable, this is defined by a Probability Mass Function (PMF), whereas for continuous variables, it relies on a Probability Density Function (PDF).

Common examples include:

  • Normal (Gaussian) Distribution: Symmetric and bell-shaped, completely described by its mean and variance. It is foundational in statistics due to the Central Limit Theorem.
  • Binomial Distribution: Models the number of successes in a fixed number of independent Bernoulli trials (e.g., coin flips).
  • Poisson Distribution: Models the number of events occurring in a fixed interval of time or space, assuming events happen at a constant average rate independently (e.g., emails arriving per hour).

“Explain Singular Value Decomposition and its applications to ML.”

Quick answer

SVD factorizes a matrix into three matrices: U, Sigma, and V-transpose. It represents data in terms of its principal orthogonal components. It is widely used in PCA, recommendation systems, and data compression.

💡 Note SVD factorizes a matrix into three matrices: U, Sigma, and V-transpose. It represents data in terms of its principal orthogonal components. It is widely used in PCA, recommendation systems, and data compression.

Answer

Singular Value Decomposition (SVD) is a cornerstone of linear algebra that factorizes any m×nmatrixm \times n matrix Aintothreematrices:into three matrices:A = U \Sigma V^T$.

  • UU is an m×mm \times m orthogonal matrix (left singular vectors).
  • Σ\Sigma is an m×nm \times n diagonal matrix containing singular values (scaling factors, sorted in descending order).
  • VTV^T is an n×nn \times n orthogonal matrix (right singular vectors).

SVD essentially says that any linear transformation can be broken down into a rotation (VTV^T), a scaling along coordinate axes (Σ\Sigma), and another rotation (UU).

In machine learning, SVD is incredibly powerful:

  1. PCA: SVD is the numerically stable way to compute Principal Component Analysis without forming the covariance matrix directly.
  2. Dimensionality Reduction & Compression: Truncating Σ\Sigma to keep only the largest singular values provides the best low-rank approximation of the original data.
  3. Recommendation Systems: Matrix factorization algorithms (like those used in the Netflix prize) use variations of SVD to uncover latent features connecting users and items.

“Explain Type I and Type II errors and statistical power.”

Quick answer

Type I error is a false positive (rejecting a true null hypothesis). Type II error is a false negative (failing to reject a false null). Power is the probability of correctly rejecting a false null.

💡 Note Type I error is a false positive (rejecting a true null hypothesis). Type II error is a false negative (failing to reject a false null). Power is the probability of correctly rejecting a false null.

Answer

In hypothesis testing, we make decisions about a null hypothesis (H0H_0), leading to two possible types of errors:

  • Type I Error (False Positive): Rejecting H0H_0 when it is actually true. The probability of this is denoted by α\alpha (the significance level).
  • Type II Error (False Negative): Failing to reject H0H_0 when the alternative hypothesis is true. The probability of this is denoted by β\beta.

Statistical Power is 1−β1 - \beta. It represents the probability of correctly rejecting the null hypothesis when there is a true effect. High power means a test is highly sensitive to detecting an effect if one exists.

In machine learning and A/B testing, minimizing Type I error ensures we don't adopt useless features or UI changes, while maximizing Power ensures we don't miss out on actual improvements. Increasing sample size is the most common way to increase power without sacrificing α\alpha.

“What is variance and standard deviation?”

Quick answer

Variance measures the average squared deviation from the mean, while standard deviation is the square root of variance, keeping the same units as the original data.

💡 Note Variance measures the average squared deviation from the mean, while standard deviation is the square root of variance, keeping the same units as the original data.

Answer

Variance quantifies how much the data points in a distribution spread out from the mean. It is calculated as the average of the squared differences from the mean. Because the differences are squared, variance gives more weight to extreme values (outliers) and its units are squared compared to the original data.

Standard Deviation is simply the square root of the variance. This brings the measure of spread back to the original units of the data, making it much easier to interpret. In ML, standard deviation is heavily used in standardizing features (e.g., Z-score normalization) to ensure that models like SVMs and neural networks treat all features equally regardless of their original scale.

“What is a vector and what is a dot product?”

Quick answer

A vector is an array of numbers representing a point or direction in space. A dot product is an algebraic operation that takes two equal-length sequences of numbers and returns a single scalar.

💡 Note A vector is an array of numbers representing a point or direction in space. A dot product is an algebraic operation that takes two equal-length sequences of numbers and returns a single scalar.

Answer

A vector is a mathematical entity that has both magnitude and direction, commonly represented as an ordered array of numbers in Rn\mathbb{R}^n. In ML, vectors usually represent data points (features) or weights.

The dot product (or scalar product) is an algebraic operation between two vectors of the same size. For vectors a\mathbf{a} and b\mathbf{b}, it is calculated as a⋅b=∑i=1naibi\mathbf{a} \cdot \mathbf{b} = \sum_{i=1}^{n} a_i b_i.

Geometrically, the dot product is related to the angle θ\theta between the vectors: a⋅b=∥a∥∥b∥cos⁡(θ)\mathbf{a} \cdot \mathbf{b} = \|\mathbf{a}\| \|\mathbf{b}\| \cos(\theta). If the dot product is 0, the vectors are orthogonal (perpendicular). In ML, the dot product is heavily used in computing projections, similarity measures (like cosine similarity), and the linear combinations in neural network layers.