Skip to content
AI360Xpert
Beta
LLM Fine-Tuning & Alignment

LLM Fine-Tuning & Alignment

30 interview questions in this topic, each with its full answer shown below. Use "Collapse all" to skim just the titles.

“What is the difference between a base model and a chat model?”

Quick answer

A base model predicts the next word based on raw text data, while a chat model has been fine-tuned on conversational data to follow instructions and interact naturally with users.

A chat model (or instruct model) starts as a base model but undergoes significant post-training to behave like a helpful assistant. This involves Supervised Fine-Tuning (SFT) on question-answer pairs and often Reinforcement Learning from Human Feedback (RLHF) to align its responses with human preferences.

Because of this tuning, a chat model understands that it should directly answer prompts, maintain a conversational tone, refuse harmful requests, and follow complex formatting instructions.

💡 Note While base models are rarely used directly by end-users, they are excellent starting points for developers who want to perform their own specialized fine-tuning for domain-specific applications.

“What is catastrophic forgetting, and how do you reduce it during fine-tuning?”

Quick answer

Catastrophic forgetting occurs when a model forgets its pre-trained general knowledge while learning a new specific task. It can be reduced using low learning rates, PEFT methods like LoRA, or data mixing.

This happens because the weight updates during fine-tuning overwrite the representations that encoded the original knowledge.

Strategies to mitigate catastrophic forgetting include:

  • Parameter-Efficient Fine-Tuning (PEFT): Methods like LoRA freeze the base model weights entirely, training only a tiny set of adapter weights. Because the base model is untouched, its core knowledge remains completely intact.
  • Data Mixing / Replay: When fine-tuning, include a small percentage (e.g., 5-10%) of the original pretraining or general instruction data alongside the new task-specific data. This forces the model to maintain its general capabilities.
  • Low Learning Rates and Early Stopping: Use very conservative learning rates and stop training the moment the model performs adequately on the new task, preventing drastic weight shifts.

💡 Note If your specific task strictly limits the model's domain (e.g., a customer service bot that should only talk about refunds), mild catastrophic forgetting of general knowledge might actually be desirable.

“How do you extend a model's context length, for example with RoPE scaling or sliding-window attention, and what degrades?”

Quick answer

Context length is extended using RoPE scaling to interpolate position embeddings for longer sequences, or Sliding Window Attention to restrict memory usage. Extending context often degrades the model's ability to recall information perfectly, especially in the middle of long texts.

RoPE (Rotary Position Embedding) Scaling: Models use positional embeddings to understand word order. If a model was trained on 4k tokens, it doesn't know what position 5,000 looks like. RoPE scaling (like linear or dynamic NTK-aware scaling) works by mathematically compressing the positional representations. Instead of extrapolating to unseen positions, it squishes the new, longer sequence into the original 0-4k positional space. The model is then fine-tuned briefly to adapt to this "compressed" spatial awareness.

Sliding Window Attention (SWA): Standard attention has quadratic memory cost. SWA modifies the attention mechanism so a token only looks at a fixed window of recent tokens (e.g., the last 4k tokens), rather than the entire history. This creates a linear memory requirement, allowing near-infinite context, but compromises long-range dependencies.

What Degrades? Extending context reliably introduces the "Lost in the Middle" phenomenon. While the model remembers the very beginning and very end of a massive prompt perfectly, its accuracy and recall for facts buried in the middle of the expanded context window degrade significantly.

💡 Note Before deploying context extension techniques, consider if a robust RAG (Retrieval-Augmented Generation) system would better serve the use case by only feeding the most relevant snippets to the model.

“What is continuous batching, and why does it improve throughput?”

Quick answer

Continuous batching dynamically inserts new requests into an active batch at the token level, eliminating wait times for shorter requests to finish and drastically increasing GPU utilization and throughput.

In naive static batching, multiple requests are grouped together to process concurrently. Because LLM generation is variable in length, the entire batch must wait until the longest request finishes generating all its tokens. Shorter requests that finish early sit idle, wasting expensive GPU compute cycles.

Continuous batching operates at the iteration (token) level rather than the request level. When a request within the batch finishes and emits an end-of-sequence token, it is immediately ejected from the batch. The server then instantly pulls a new request from the queue and inserts it into the newly freed slot for the very next token generation step.

By ensuring the GPU is always working on a full batch of tokens, continuous batching eliminates idle compute time, often resulting in throughput improvements of 10x or more for diverse workloads.

💡 Note Continuous batching requires complex memory management for the KV cache, which is why it is almost always paired with PagedAttention to prevent memory fragmentation as requests enter and exit the batch.

“Compare greedy decoding, beam search, top-k, and top-p sampling.”

Quick answer

Greedy decoding picks the most likely token; beam search tracks multiple sequences for better overall likelihood; top-k samples from the K most likely tokens; top-p samples from a dynamic pool representing cumulative probability p.

Greedy Decoding is the simplest method: it consistently selects the single token with the highest probability. While fast and deterministic, it can lead to repetitive or sub-optimal text because it doesn't consider future word sequences.

Beam Search explores multiple possible paths. It maintains the top BB (beam width) most likely sequences at each step. By looking ahead, it avoids greedy decoding's local traps, producing higher quality, coherent text, though it is computationally expensive and less creative.

Top-k Sampling introduces randomness by restricting the model to sample only from the kk most probable next tokens. It cuts off the "long tail" of highly unlikely words, preventing gibberish while maintaining variety. However, a fixed kk can be restrictive if there are many good options, or too broad if there is only one logical choice.

Top-p Sampling (Nucleus Sampling) dynamically adjusts the pool of tokens. It samples from the smallest set of top tokens whose cumulative probability exceeds a threshold pp. This ensures that in confident situations, it samples from a few tokens, and in uncertain situations, it samples from many, offering better contextual generation than Top-k.

💡 Note Top-p and Top-k are often used together in modern LLM APIs to fine-tune the balance between coherence and creativity.

“Compare DPO, PPO-based RLHF, and RLAIF for alignment. What are the trade-offs?”

Quick answer

PPO uses a separate reward model for complex alignment but is unstable. DPO directly optimizes the model on preference data, making it simpler and more stable. RLAIF uses an LLM instead of humans to generate preference data, saving time and money.

PPO-based RLHF (Proximal Policy Optimization): The traditional method. It trains a separate Reward Model (RM) on human preference data, then uses the PPO reinforcement learning algorithm to optimize the LLM against the RM.

  • Pros: Highly effective; allows for complex, dynamic exploration during training.
  • Cons: Extremely complex to implement, highly unstable, requires maintaining multiple large models in memory simultaneously (Actor, Reference, Reward, Critic).

DPO (Direct Preference Optimization): DPO mathematically bypasses the need for a separate Reward Model and reinforcement learning. It directly updates the LLM's weights using a cross-entropy loss that increases the probability of the preferred response and decreases the rejected one.

  • Pros: Much simpler, mathematically stable, requires fewer resources (only the target and reference model).
  • Cons: Prone to overfitting on the specific preference dataset; struggles slightly with out-of-distribution generalization compared to PPO.

RLAIF (Reinforcement Learning from AI Feedback): Instead of paying humans to rank responses (which is slow and expensive), RLAIF uses a powerful, pre-aligned LLM (like GPT-4) to act as the judge and generate the preference data.

  • Pros: Massively scalable, cheap, and fast.
  • Cons: Inherits the biases and limitations of the AI judge; can lead to a "mode collapse" towards the AI judge's specific style.

💡 Note Currently, DPO is the industry standard for open-weights fine-tuning due to its balance of simplicity and performance, effectively replacing PPO for many organizations.

“How do you evaluate a fine-tuned model against the base model?”

Quick answer

Evaluation requires quantitative benchmarks for domain accuracy, qualitative human or LLM-as-a-judge reviews for tone and alignment, and regression testing to ensure general capabilities weren't lost.

A robust evaluation pipeline involves several layers:

  1. Task-Specific Quantitative Metrics: Evaluate the model on a holdout test set that mirrors the fine-tuning data. If the task is classification or structured extraction, use traditional metrics like precision, recall, and F1 score. For summarization, use ROUGE or BLEU, though they have limitations in semantic understanding.
  2. LLM-as-a-Judge: Because traditional NLP metrics struggle with open-ended generation, use a superior model (like GPT-4) to evaluate the fine-tuned model's outputs. You provide a rubric, and the "judge" model scores the output on helpfulness, accuracy, and adherence to constraints.
  3. Human Evaluation: For final alignment, blinded A/B testing with human domain experts is essential to catch subtle nuances, tone issues, or domain-specific hallucinations that automated systems miss.
  4. Regression Benchmarks: Run the fine-tuned model against standard LLM benchmarks (like MMLU, HumanEval, or GSM8K) and compare the scores to the base model to quantify how much general knowledge was compromised.

💡 Note Always track inference latency during evaluation. If a fine-tuned model requires a massive system prompt to function properly, it may be too slow or expensive for production use compared to the base model.

“What is fine-tuning, and how does it differ from pretraining?”

Quick answer

Pretraining teaches a model general language representation using vast unlabelled text, whereas fine-tuning adapts the pre-trained model to specific tasks using a smaller, task-specific dataset.

Fine-tuning takes this pre-trained base model and continues training it on a much smaller, curated dataset tailored to a specific task or domain, such as summarization, sentiment analysis, or coding. Unlike pretraining, fine-tuning updates the model's weights to optimize for the specific prompt-response pairs it is shown. Because the base model already understands language, fine-tuning requires significantly less compute and data.

💡 Note Fine-tuning does not inject a massive amount of new factual knowledge; rather, it teaches the model how to format its output and behave in a specific manner based on its existing knowledge.

“Explain FlashAttention and why it is faster and more memory-efficient than standard attention.”

Quick answer

FlashAttention is an algorithm that computes exact attention while avoiding large memory reads/writes by fusing operations and keeping intermediate matrices in the fast GPU SRAM instead of slow HBM.

FlashAttention is a mathematically exact algorithm that reorganizes the attention computation to be "hardware aware." Its primary insight is that GPU memory bandwidth (moving data between slow HBM and fast SRAM) is the actual bottleneck, not compute capability.

FlashAttention relies on two techniques: Tiling and Recomputation. It uses tiling to break the massive Q,K,VQ, K, V matrices into smaller blocks that fit perfectly into the GPU's ultra-fast, on-chip SRAM. It computes the attention for these blocks locally in SRAM, completely avoiding the need to write and read the massive intermediate N×NN \times N matrix to HBM. During the backward pass for training, it uses recomputation to calculate necessary values on the fly rather than storing them, saving massive amounts of memory.

💡 Note FlashAttention is largely responsible for the feasibility of modern massive context windows (like 128k or 1M tokens), as standard attention would instantly trigger Out-Of-Memory errors at those sizes.

“How would you decide between full fine-tuning, LoRA, and prefix or prompt tuning for a given budget and task?”

Quick answer

Use prompt tuning for simple formatting with strict budgets. Use LoRA for 95% of use cases, balancing high performance with low cost. Reserve full fine-tuning for massive domain shifts or when building foundational industry models with large budgets.

Prompt / Prefix Tuning: These methods train a small set of "virtual tokens" appended to the input, leaving the model frozen.

  • When to use: You have an extremely restricted budget, minimal GPU resources, and the task is simple (like tone adjustment or basic formatting). It scales poorly to complex tasks.

LoRA (Low-Rank Adaptation): LoRA trains small adapter matrices injected into the model layers, keeping the base model frozen.

  • When to use: This is the industry default for almost all fine-tuning tasks. It requires minimal VRAM (can be done on single GPUs), trains quickly, and achieves 95-99% of the performance of full fine-tuning. It is ideal for instruction tuning, teaching JSON schemas, or domain adaptation.

Full Fine-Tuning: Updates every single parameter in the massive neural network.

  • When to use: You have a massive budget, a huge dataset, and you are introducing entirely new languages, fundamental domain shifts (e.g., from general text to complex biological sequences), or creating a new base model. It is highly susceptible to catastrophic forgetting and requires immense multi-node GPU clusters.

💡 Note Always start with Prompt Engineering. If that fails, move to LoRA. Only consider full fine-tuning if extensive LoRA experiments hit a strict performance ceiling on a specialized task.

“What is knowledge distillation?”

Quick answer

Knowledge distillation is a compression technique where a smaller 'student' model is trained to mimic the behavior and outputs of a larger, more capable 'teacher' model.

Instead of training the student model solely on raw datasets, it is trained to replicate the teacher model's outputs. Specifically, the student attempts to match the teacher's probability distribution (the logits) for predicting the next word. Because the teacher's outputs contain "soft labels" (e.g., showing that "car" is highly likely, "truck" is somewhat likely, but "apple" is unlikely), the student learns deeper contextual relationships and generalizations that are not present in strict one-hot encoded training data.

This technique is widely used to deploy capable AI on edge devices or in low-latency environments, as the student model requires a fraction of the compute power of the teacher while retaining a significant portion of its performance.

💡 Note In the era of LLMs, distillation often takes the form of using a massive model (like GPT-4) to generate high-quality synthetic datasets that are then used to fine-tune a smaller open-source model (like Llama 3 8B).

“What is the KV cache, and why does it speed up autoregressive decoding?”

Quick answer

The KV cache stores previously computed Key and Value vectors in attention layers, preventing the LLM from redundantly recalculating them for past tokens during step-by-step text generation.

During each step of the attention mechanism, a token is projected into Query (Q), Key (K), and Value (V) vectors. To determine context, the Query of the current token must attend to the Keys and Values of all preceding tokens. Without optimization, the model would needlessly recompute the K and V vectors for every single past token at every generation step, leading to massive redundant computation and incredibly slow inference.

The KV Cache solves this by storing the computed Key and Value tensors for all previous tokens in GPU memory. When generating a new token, the model only computes the Q, K, and V for that specific new token, and retrieves the past K and V vectors from the cache.

💡 Note While the KV cache drastically reduces compute latency, it becomes the primary bottleneck for memory consumption during inference, especially with long context windows and large batch sizes.

“How would you design an evaluation and regression-testing pipeline for a fine-tuned LLM before release?”

Quick answer

Design a pipeline combining automated domain-specific benchmarks, LLM-as-a-judge for subjective quality, regression testing on standard datasets to catch capability loss, and rigorous red-teaming for security and safety.

  1. Domain-Specific Automated Testing: Create a golden dataset of highly curated prompts and ideal outputs for your specific task. Use an "LLM-as-a-judge" (e.g., GPT-4) with strict rubrics to score the fine-tuned model's responses on accuracy, tone, and formatting constraints against this golden set.
  2. Regression Benchmarks: To detect catastrophic forgetting, run the fine-tuned model through standard open-source benchmarks (e.g., MMLU for knowledge, HumanEval for coding, GSM8K for math). Compare these baseline scores to the original base model. A drop of 1-3% is acceptable; a massive drop indicates a failed fine-tuning run.
  3. Red-Teaming and Safety Evaluation: Subject the model to adversarial prompts to ensure the fine-tuning process didn't break its safety guardrails. Check for susceptibility to prompt injection, jailbreaks, and the generation of toxic or biased content.
  4. System and Performance Testing: Deploy the model to a staging environment and run load testing. Measure Time to First Token (TTFT), Time Per Output Token (TPOT), and VRAM usage. A model that is perfectly accurate but too slow or memory-hungry cannot be released.

💡 Note Version control your datasets and evaluation prompts just like code. A model evaluation is only valid if the pipeline running it is deterministic and version-locked.

“What is quantization in the context of LLMs?”

Quick answer

Quantization is a technique to reduce the memory footprint and improve inference speed of LLMs by converting high-precision floating-point weights into lower-precision formats like 8-bit or 4-bit integers.

Typically, models are trained using 32-bit (FP32) or 16-bit (FP16/BF16) floating-point numbers. Quantization maps these high-precision values to lower-precision data types, such as 8-bit (INT8) or even 4-bit (INT4) integers. By doing so, the memory required to load the model into GPU VRAM is drastically reduced—often by 2x to 4x or more. This allows massive models, like a 70B parameter model, to run on consumer hardware or fewer enterprise GPUs.

While quantization significantly improves memory efficiency and memory bandwidth utilization (which often bottlenecks LLM inference), it introduces a small amount of approximation error.

💡 Note Modern quantization schemes, such as AWQ (Activation-aware Weight Quantization) or GPTQ, strategically preserve the most critical weights in higher precision to minimize the loss in model quality and reasoning capability.

“What is a LoRA adapter at a high level?”

Quick answer

A LoRA adapter is a small, pluggable set of trained weights that modifies a frozen base model's behavior for a specific task, requiring significantly less memory and storage than full fine-tuning.

Instead of updating all billions of parameters in a model (which requires immense computational power and memory), LoRA injects small, trainable "adapter" modules into the layers of the network. During training, the base model remains completely frozen, and only the small adapter weights are updated.

Once training is complete, the LoRA adapter is saved as a separate, very small file (often just a few megabytes), compared to the multi-gigabyte base model. During inference, this adapter is loaded alongside the base model and its outputs are merged with the base model's outputs to produce task-specific behavior.

💡 Note The beauty of LoRA adapters is modularity. You can host a single massive base model in memory and dynamically swap small LoRA adapters in and out to serve dozens of different customized tasks efficiently.

“How does LoRA work, and why is it parameter-efficient? Compare it with QLoRA.”

Quick answer

LoRA uses low-rank matrix decomposition to approximate weight updates, training a tiny fraction of parameters. QLoRA pushes efficiency further by quantizing the base model to 4-bit, enabling fine-tuning on consumer GPUs.

QLoRA (Quantized LoRA) extends this efficiency by addressing the memory footprint of the frozen base model itself. While standard LoRA loads the base model in 16-bit precision, QLoRA aggressively quantizes the frozen base model down to 4-bit precision using a specialized data type (NormalFloat4).

During QLoRA training, the 4-bit weights are briefly dequantized to 16-bit to compute gradients, which are then passed to the 16-bit LoRA adapters. This breakthrough allows massively parameter-efficient fine-tuning (e.g., a 65B model) on a single high-end consumer GPU.

💡 Note QLoRA trades a small amount of training speed (due to the constant quantize/dequantize overhead) for a massive reduction in VRAM requirements.

“How do Mixture-of-Experts models work, and what are the serving challenges?”

Quick answer

MoE models replace standard layers with multiple specialized 'expert' neural networks and a router network that sends tokens to only the relevant experts. They offer massive parameter counts with fast inference, but require immense VRAM.

In a standard dense transformer, every token passes through every parameter in the Feed-Forward Network (FFN) layers. In an MoE model, the FFN is replaced by several independent sub-networks called "experts" (e.g., 8 experts). A trainable "router" network analyzes each incoming token and directs it to only the top KK experts (usually K=2K=2).

This means a model might have 47 Billion total parameters (like Mixtral 8x7B), but only ~13 Billion parameters are "active" for any given token. This provides the reasoning power of a massive model with the inference speed of a much smaller one.

Serving Challenges: While computationally cheap, MoE models are a nightmare for memory management.

  • VRAM Footprint: You still must load all 47B parameters into GPU memory, requiring high-end multi-GPU setups just to host the model, even though compute utilization is low.
  • Load Balancing: The router might favor one or two experts heavily (e.g., all coding questions go to Expert 3). This causes a bottleneck where one GPU is overworked while others sit idle. Advanced serving engines must implement complex token dropping or expert parallelism to manage these imbalances.

💡 Note MoE is the architecture utilized by many frontier models, including GPT-4, as it allows for scaling up intelligence while keeping inference latency manageable.

“Compare PagedAttention-style memory management with naive KV cache allocation.”

Quick answer

Naive KV cache statically allocates maximum memory for requests, causing massive fragmentation. PagedAttention divides memory into fixed blocks, allocating them dynamically like virtual memory, virtually eliminating fragmentation and enabling higher batch sizes.

In naive KV cache allocation, the system must allocate a contiguous block of GPU memory for the maximum possible length of a request as soon as it arrives. Because the final output length is unpredictable, much of this allocated memory goes unused. Furthermore, as requests finish at different times, the memory becomes heavily fragmented, making it impossible to fit new requests even if total free memory is high. This limits maximum batch size severely.

PagedAttention (introduced by vLLM) solves this by borrowing the concept of virtual memory and paging from operating systems. It divides the KV cache into fixed-size blocks (pages), where each block holds the vectors for a small number of tokens (e.g., 16 tokens).

Instead of contiguous allocation, PagedAttention maps continuous logical tokens to non-contiguous physical blocks in GPU memory. Blocks are allocated dynamically only when they are needed. This eliminates internal fragmentation and sharing of blocks between identical prompt prefixes becomes possible, allowing the system to serve significantly larger batches and increase throughput.

💡 Note PagedAttention is the foundational technology that makes continuous batching highly effective in modern production LLM serving architectures.

“What is prompt caching, and when does it save cost?”

Quick answer

Prompt caching stores the computed KV cache of frequently used prompt prefixes, saving computation time and cost by bypassing the heavy initial processing phase for repeated context.

When a model processing a prompt, it computes the Key and Value (KV) vectors for every token during the pre-fill phase. If users send prompts that share a large, identical prefix (e.g., a massive system prompt, a large document in a RAG pipeline, or few-shot examples), recalculating the KV vectors for that shared prefix every single time wastes immense compute.

Prompt caching stores the KV cache of these common prefixes in memory (or disk). When a new request arrives, the system checks if its prefix matches a cached sequence. If it does, the model simply loads the cached KV vectors and only processes the new, unique tokens appended to the end.

This dramatically reduces "Time to First Token" (TTFT) and saves significant compute costs, as APIs often charge drastically less for cached input tokens.

💡 Note To leverage prompt caching effectively in commercial APIs, developers must ensure the static, shared context is always placed at the very beginning of the prompt, as caching operates strictly on exact prefix matches.

“When would you choose prompting, RAG, or fine-tuning?”

Quick answer

Use prompting for quick experiments, RAG for tasks requiring up-to-date or external factual knowledge, and fine-tuning to adapt model behavior, tone, or handle highly specialized domain structures.

Prompting (or Prompt Engineering) is the easiest and cheapest method. It involves carefully crafting the input text to guide the model's response. It is ideal for general tasks, rapid prototyping, and when the model already possesses the required knowledge and capabilities.

RAG (Retrieval-Augmented Generation) is best when the model needs access to dynamic, proprietary, or highly specific factual information that wasn't in its training data. By retrieving relevant documents and injecting them into the prompt, RAG grounds the model in facts and significantly reduces hallucination, without needing to retrain the model.

Fine-tuning is necessary when you need to change the model's behavior, tone, or format, rather than just injecting new facts. It is ideal for specialized domains (like legal or medical text parsing), teaching a specific JSON output schema, or when prompt context limits are a bottleneck.

💡 Note These techniques are not mutually exclusive. A robust production system often uses a fine-tuned model combined with RAG to achieve the best performance and accuracy.

“Compare post-training quantization and quantization-aware training. What is the accuracy trade-off?”

Quick answer

PTQ quantizes a fully trained model, which is fast but can cause accuracy drops. QAT simulates quantization during training, allowing the model to adapt to precision loss, resulting in better accuracy at the cost of training time.

Post-Training Quantization (PTQ) is applied after the model has been fully trained in high precision. It involves mapping the existing weights to lower precision (like 8-bit or 4-bit) using a calibration dataset. PTQ is extremely fast and computationally cheap because it requires no backpropagation or retraining. However, because the weights are "forced" into lower precision without the model adapting to it, aggressive PTQ can lead to significant accuracy degradation and increased hallucination.

Quantization-Aware Training (QAT) introduces quantization during the training or fine-tuning phase itself. The "forward pass" simulates the low-precision quantization, and the model calculates the loss. The "backward pass" updates the high-precision weights based on this simulated error. Because the model is actively learning to compensate for the rounding errors introduced by quantization, QAT yields significantly higher accuracy and preserves reasoning capabilities much better than PTQ.

💡 Note For extremely large LLMs (like 70B+ parameters), full QAT is often prohibitively expensive. In such cases, advanced PTQ methods like AWQ or GPTQ are preferred as they intelligently minimize the accuracy drop without retraining.

“How would you detect and mitigate reward hacking when training with a learned reward model?”

Quick answer

Detect reward hacking by monitoring for high reward scores combined with degraded human evaluation or linguistic diversity. Mitigate it using KL-divergence penalties and regularizing the reward model.

Detection:

  • Divergence Metrics: Monitor the linguistic diversity (e.g., perplexity, vocabulary usage) of the generated text. A sudden drop indicates the model is exploiting a specific phrase.
  • Human-in-the-Loop: Periodically sample generations that receive high RM scores and have human annotators verify them. A disconnect between the RM score and human judgment confirms hacking.

Mitigation:

  • KL-Divergence Penalty: This is the most crucial defense. During PPO, you penalize the LLM if its output probability distribution diverges too far from the original, safe Supervised Fine-Tuned (SFT) model. This anchors the model to human-like text.
  • Ensemble Reward Models: Train multiple reward models and take their average score, or use the lowest score among them, making it harder for the LLM to exploit a single model's blind spots.

💡 Note Reward hacking is fundamentally an issue of out-of-distribution exploitation. The LLM enters a state space the RM was never trained on, rendering the RM's judgments meaningless.

“What is RLHF, and what are its three main stages?”

Quick answer

RLHF aligns LLMs with human preferences. Its stages are supervised fine-tuning (SFT), training a reward model on human-ranked responses, and optimizing the LLM using reinforcement learning (like PPO).

The RLHF pipeline consists of three distinct stages:

  1. Supervised Fine-Tuning (SFT): The base model is fine-tuned on a high-quality dataset of human-written prompts and ideal responses. This teaches the model the basic format of dialogue and instruction-following, resulting in an SFT model.
  2. Training a Reward Model (RM): The SFT model generates multiple different responses to a set of prompts. Human annotators rank these responses from best to worst based on helpfulness, safety, and accuracy. A separate model (the Reward Model) is trained on these rankings to predict a scalar score representing human preference.
  3. Reinforcement Learning (RL) Optimization: The SFT model is further trained using a reinforcement learning algorithm, most commonly Proximal Policy Optimization (PPO). The model generates responses, the Reward Model scores them, and PPO updates the LLM's weights to maximize the reward while penalizing it for drifting too far from the original SFT model's behavior.

💡 Note The KL-divergence penalty in the PPO stage is critical; without it, the model might "hack" the reward model, producing gibberish that statistically yields high scores.

“How would you serve a 70B-parameter model at low latency across several GPUs? Discuss tensor parallelism and quantization.”

Quick answer

Serving a 70B model requires Tensor Parallelism to split weights across multiple GPUs to pool VRAM and compute, combined with 8-bit or 4-bit quantization to reduce memory overhead and increase processing speed.

To achieve low latency, two primary strategies are required:

1. Tensor Parallelism (TP): Instead of placing the whole model on one GPU, Tensor Parallelism shards the individual matrix multiplication operations across multiple GPUs (e.g., 2, 4, or 8 GPUs) linked by a high-speed interconnect like NVLink. Each GPU holds a fraction of the weights and computes a fraction of the mathematical operations simultaneously. They then synchronize their results. This pools the VRAM and drastically reduces the compute time per token, achieving low latency.

2. Quantization: To further reduce latency and increase the batch size capability, apply quantization (e.g., FP8, INT8, or AWQ/GPTQ INT4). By reducing the weights to 8-bit, the memory footprint halves to ~70GB. This means the model requires fewer GPUs to host, and significantly less time is spent transferring weights from memory to the compute cores (the primary bottleneck in generation), drastically lowering Time Per Output Token.

💡 Note When combining these, you could use an inference engine like vLLM or TensorRT-LLM, employing TP=4 and FP8 quantization, alongside PagedAttention for maximum throughput.

“How do you prepare a high-quality dataset for supervised fine-tuning?”

Quick answer

Preparation involves defining the target format, gathering diverse and high-quality prompt-completion pairs, deduplicating, removing toxic/biased data, and strictly ensuring data accuracy and consistency.

Preparing an SFT dataset involves several rigorous steps:

  1. Curation and Generation: Collect prompt-response pairs that perfectly match the desired behavior, tone, and formatting. This can be done via human experts (highly accurate but expensive) or synthesized using larger, capable models like GPT-4 (faster, but requires verification).
  2. Diversity and Distribution: Ensure the dataset covers a wide variety of topics, prompt lengths, and complexities. If training an assistant, include edge cases like refusals, clarifying questions, and logic puzzles.
  3. Cleansing and Formatting: Remove spelling errors, formatting inconsistencies, and factual inaccuracies. Standardize the structural formatting (e.g., using specific <user> and <assistant> tags).
  4. Deduplication and Filtering: Remove duplicate entries to prevent the model from memorizing specific answers. Aggressively filter out toxic, biased, or PII (Personally Identifiable Information) contaminated data.

💡 Note A common practice is to train on a small, hyper-curated subset of the data first. If the model fails to learn the behavior on this "golden dataset," the issue is likely the training parameters, not dataset size.

“Explain speculative decoding and the conditions under which it helps.”

Quick answer

Speculative decoding uses a small, fast draft model to predict multiple tokens, which a large target model verifies in parallel. It drastically speeds up inference when memory bandwidth, not compute, is the bottleneck.

Because LLM inference is memory-bandwidth bound (reading the massive model weights into the processor for every single token takes longer than the actual math), generating one token at a time is inefficient.

Speculative decoding introduces a smaller, much faster "draft" model.

  1. The small draft model rapidly predicts a sequence of future tokens (e.g., 4 tokens).
  2. The large, accurate "target" model takes this sequence and processes all 4 tokens in a single parallel step to verify them.
  3. If the target model agrees with the draft tokens, all 4 are accepted, effectively yielding 4 tokens in the time it usually takes to generate 1. If it disagrees at token 2, it rejects the rest and generates the correct token 2 itself.

It helps most under low batch-size conditions (where compute is underutilized) and when generating highly predictable text (like code syntax or common phrases), ensuring high acceptance rates from the draft model.

💡 Note Speculative decoding guarantees that the final output is mathematically identical to what the large target model would have produced alone, making it a "free" speedup if the hardware setup supports it.

“What is temperature in text generation?”

Quick answer

Temperature is a hyperparameter that controls the randomness of an LLM's output. A low temperature makes the output deterministic and focused, while a high temperature increases creativity and variance.

When an LLM predicts the next word, it outputs a probability distribution over the entire vocabulary. Temperature scales the logits (the raw scores before converting to probabilities) prior to applying the softmax function.

  • Low Temperature (e.g., 0.1 - 0.3): The model becomes more confident in its top choices, resulting in highly deterministic, predictable, and focused text. This is ideal for tasks requiring factual accuracy, coding, or data extraction.
  • High Temperature (e.g., 0.7 - 1.0+): The probability distribution is flattened, making lower-probability words more likely to be selected. This increases the creativity, variety, and unpredictability of the output, which is useful for brainstorming, story writing, or generating diverse responses.

💡 Note A temperature of 0 effectively turns sampling into greedy decoding, where the model consistently picks the single most probable next token without any randomness.

“What are tokens, and why do they affect cost and latency?”

Quick answer

Tokens are the fundamental units of text (subwords or characters) that LLMs process. They dictate cost and latency because models compute and bill linearly or quadratically per token generated and processed.

Tokens directly impact both cost and latency in LLM systems. Commercial APIs (like OpenAI's or Anthropic's) charge per 1,000 or 1,000,000 tokens processed. Therefore, longer prompts and longer generated outputs inherently cost more.

Regarding latency, generating text is an autoregressive process; the model must compute and generate each token sequentially. The time it takes to generate a response (Time Per Output Token, TPOT) scales linearly with the number of generated tokens. Additionally, processing a massive number of input tokens increases the "Time to First Token" (TTFT) due to the computation required to read and encode the prompt.

💡 Note A common rule of thumb for English text is that 1 token is approximately equal to ¾ of a word, though this varies significantly for code or non-English languages.

“What is a context window?”

Quick answer

A context window is the maximum number of tokens an LLM can process in a single pass, dictating how much previous text it can 'remember' and consider when generating the next word.

Because of the self-attention mechanism in transformer architectures, the memory and computational requirements traditionally scale quadratically with the sequence length. Thus, if a model has a context window of 8,192 tokens, any input exceeding this limit must be truncated, or the model will fail to process it. A larger context window allows a model to ingest entire books, codebases, or lengthy conversation histories, providing better contextual understanding and more coherent long-form generation.

💡 Note When building applications, exceeding the context window usually results in losing the earliest parts of the conversation. Techniques like RAG or sliding-window memory are often used to manage long interactions within a fixed context limit.

“What is instruction tuning?”

Quick answer

Instruction tuning is a form of fine-tuning where an LLM is trained on examples of instructions and their corresponding correct responses to improve its ability to follow user commands.

During instruction tuning, the model is trained on a dataset comprising pairs of instructions (prompts) and their desired outputs (responses). For example, a prompt might be "Translate the following sentence to French: 'Hello world'", and the target response would be "Bonjour le monde". By updating the model's weights to minimize the difference between its predictions and these targets, the model learns the concept of instruction-following across various tasks.

This process is critical for creating modern chat models and assistants. It bridges the gap between raw text generation capabilities and practical usability, making the model much more aligned with human intent.

💡 Note Instruction tuning datasets often contain a wide variety of tasks (e.g., summarizing, coding, writing poetry) to ensure the model generalizes well to unseen instructions, a property known as zero-shot generalization.