Paper Breakdowns
Plain-language walkthroughs of the papers that shaped modern AI — the core idea, the results that mattered, and how it connects to what you already know.
Attention Is All You Need
Attention Is All You Need
The landmark 2017 paper that introduced the Transformer architecture, replacing RNNs with self-attention and launching the modern era of large language models.
Read breakdownLlama 3
The Llama 3 Herd of Models
The 2024 Meta technical report detailing a massive scale-up in data and compute, proving that dense models can reach frontier-level capabilities through sheer scale and data quality.
Read breakdownLlama 2
Llama 2: Open Foundation and Fine-Tuned Chat Models
The 2023 paper from Meta that introduced a commercially viable, chat-tuned foundation model, detailing the immense effort required for RLHF and safety alignment.
Read breakdownLLaMA 1
LLaMA: Open and Efficient Foundation Language Models
The 2023 paper from Meta that introduced the first highly capable, open-weights foundation model, sparking the open-source LLM revolution.
Read breakdownGPT-4
GPT-4 Technical Report
The 2023 technical report from OpenAI detailing the model that defined the frontier of AI capabilities, demonstrating human-level performance on professional benchmarks.
Read breakdownClaude 3
The Claude 3 Model Family: Opus, Sonnet, Haiku
The 2024 Anthropic paper detailing a family of models that pushed the frontier of AI capabilities, heavily utilizing Constitutional AI and Constitutional alignment.
Read breakdownGemini
Gemini: A Family of Highly Capable Multimodal Models
The 2023 Google DeepMind paper introducing a natively multimodal model family built from the ground up to reason seamlessly across text, images, audio, and video.
Read breakdownDeepSeek-R1
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
The 2025 breakthrough that proved 'aha' moments and advanced reasoning capabilities could emerge purely from large-scale reinforcement learning, without needing millions of human-written reasoning examples.
Read breakdownDeepSeek-V2
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
The 2024 paper that disrupted the AI pricing landscape by introducing Multi-Head Latent Attention (MLA), drastically reducing the memory required for the KV Cache.
Read breakdownDeepSeekMath (GRPO)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
The 2024 paper that introduced Group Relative Policy Optimization (GRPO), a highly efficient alternative to PPO that eliminates the need for a separate value model during reinforcement learning.
Read breakdownRetrieval-Augmented Generation (RAG)
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
The 2020 paper that introduced a general-purpose fine-tuning recipe for combining pre-trained parametric and non-parametric memory.
Read breakdownLoRA
LoRA: Low-Rank Adaptation of Large Language Models
The 2021 paper from Microsoft that introduced Low-Rank Adaptation, allowing massive models to be fine-tuned quickly and cheaply by only training a tiny fraction of the parameters.
Read breakdownQLoRA
QLoRA: Efficient Finetuning of Quantized LLMs
A 2023 breakthrough that combined 4-bit quantization with LoRA, making it possible to fine-tune massive 65-billion parameter models on a single consumer GPU.
Read breakdownFlashAttention-2
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
An optimized iteration of FlashAttention that better partitions work across GPU thread blocks, reaching closer to the theoretical maximum speed of the hardware.
Read breakdownFlashAttention
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
The 2022 Stanford paper that rewrote the attention algorithm to be hardware-aware, drastically speeding up Transformers and unlocking massive context windows.
Read breakdownDirect Preference Optimization (DPO)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
The 2023 Stanford paper that revolutionized LLM alignment by mathematically proving you can skip the complex Reward Model and PPO phases of RLHF, aligning models directly from preference data.
Read breakdownConstitutional AI
Constitutional AI: Harmlessness from AI Feedback
The 2022 Anthropic paper that introduced a method to align models using AI-generated feedback rather than human labels, scaling alignment securely and transparently.
Read breakdownTraining LMs to Follow Instructions with Human Feedback
Training language models to follow instructions with human feedback
The 2022 paper from OpenAI that popularized RLHF (Reinforcement Learning from Human Feedback), creating models like InstructGPT and paving the way for ChatGPT.
Read breakdownChain-of-Thought Prompting
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
The 2022 Google Brain paper that unlocked complex reasoning in LLMs by simply prompting them to show their work step-by-step before answering.
Read breakdownReAct
ReAct: Synergizing Reasoning and Acting in Language Models
A 2022 framework that interleaves reasoning (Chain-of-Thought) with acting (using external tools like Wikipedia or APIs), forming the basis of modern LLM agents.
Read breakdownTree of Thoughts
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
A 2023 paper that framed LLM reasoning as a search algorithm over a tree of possible thoughts, allowing models to evaluate their own progress, backtrack, and look ahead.
Read breakdownToolformer
Toolformer: Language Models Can Teach Themselves to Use Tools
A 2023 Meta paper that taught models to use external tools by automatically generating their own training data, fine-tuning the model to natively call APIs via special text tokens.
Read breakdownMamba
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Introduced a new Selective State Space Model architecture that achieves Transformer-level quality with linear-time inference and hardware-aware scaling.
Read breakdownMixtral 8x7B
Mixtral of Experts
The 2024 paper from Mistral AI that brought Sparse Mixture-of-Experts (MoE) to the open-source community, enabling a 47B parameter model to run at the speed of a 14B model.
Read breakdownAWQ
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Introduced Activation-aware Weight Quantization, a method that preserves LLM performance by identifying and protecting a tiny fraction of highly salient weights based on activation magnitudes.
Read breakdownGPTQ
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Introduced a highly efficient post-training quantization method that can compress 175B parameter models down to 3 or 4 bits per weight with negligible accuracy degradation.
Read breakdownBitNet (1-bit LLMs)
BitNet: Scaling 1-bit Transformers for Large Language Models
The 2023 Microsoft paper that challenged the fundamental math of neural networks, proving you can train LLMs where weights are just +1 or -1, eliminating matrix multiplication entirely.
Read breakdownEfficient Memory Management for LLM Serving
Efficient Memory Management for Large Language Model Serving with PagedAttention
Introduced PagedAttention and the vLLM engine, applying operating system paging concepts to the KV cache to eliminate memory fragmentation and drastically increase serving throughput.
Read breakdownEfficient Streaming LMs with Attention Sinks
Efficient Streaming Language Models with Attention Sinks
Discovered that LLMs naturally use the first few tokens as 'attention sinks' to dump excess attention scores, enabling a simple fix to keep models generating text infinitely without crashing.
Read breakdownFast Inference via Speculative Decoding
Fast Inference from Transformers via Speculative Decoding
Introduced Speculative Decoding, a technique that drastically speeds up LLM inference by using a small, fast draft model to generate tokens, which a larger model then verifies in parallel.
Read breakdownBetter & Faster LLMs via Multi-token Prediction
Better & Faster Large Language Models via Multi-token Prediction
Proposed training LLMs to predict the next $N$ tokens simultaneously rather than just the single next token, improving reasoning capabilities and creating a built-in draft model for fast speculative decoding.
Read breakdownLost in the Middle
Lost in the Middle: How Language Models Use Long Contexts
A landmark empirical study revealing that despite massive context windows, LLMs systematically fail to retrieve information hidden in the middle of long documents.
Read breakdownEmergent Abilities of LLMs
Emergent Abilities of Large Language Models
Argued that certain complex capabilities in LLMs appear suddenly and unpredictably only after the model crosses a specific scale threshold, becoming a central tenet of the AI scaling hypothesis.
Read breakdownAre Emergent Abilities a Mirage?
Are Emergent Abilities of Large Language Models a Mirage?
A powerful rebuttal arguing that the 'sudden' appearance of capabilities in LLMs is merely a statistical illusion caused by researchers choosing non-linear, discontinuous metrics.
Read breakdownScaling Laws for Neural Language Models
Scaling Laws for Neural Language Models
Demonstrated that language model loss decreases predictably as a power law with compute, dataset size, and parameter count, establishing the mathematical foundation for the LLM scaling race.
Read breakdownTraining Compute-Optimal LLMs
Training Compute-Optimal Large Language Models
The 'Chinchilla paper' that proved models should be scaled equally with training data, overturning the previous consensus that large models could be trained on relatively little data.
Read breakdownPhi-1
Textbooks Are All You Need
The 2023 Microsoft paper that proved massive parameter counts aren't necessary if the training data is of extremely high, 'textbook' quality.
Read breakdownOrca
Orca: Progressive Learning from Complex Explanation Traces of GPT-4
The 2023 Microsoft paper that improved upon synthetic data distillation by teaching the smaller model the *reasoning process* of the teacher model, not just the final answer.
Read breakdownSelf-Instruct
Self-Instruct: Aligning Language Models with Self-Generated Instructions
The 2022 paper that demonstrated how to use a large, proprietary LLM to generate synthetic training data to fine-tune smaller, open-source models.
Read breakdownLet's Verify Step by Step
Let's Verify Step by Step
A 2023 OpenAI paper that showed training reward models to evaluate every single step of a reasoning chain (Process Supervision) drastically outperforms models trained only to evaluate the final answer (Outcome Supervision).
Read breakdownSelf-Consistency
Self-Consistency Improves Chain of Thought Reasoning in Language Models
A 2022 paper that improved upon Chain-of-Thought prompting by generating multiple reasoning paths and taking a majority vote on the final answer.
Read breakdownLanguage Models are Few-Shot Learners
Language Models are Few-Shot Learners
The GPT-3 paper that formalized in-context learning, showing that massive scale allows models to learn new tasks simply from examples in the prompt.
Read breakdownLanguage Models are Unsupervised Multitask Learners
Language Models are Unsupervised Multitask Learners
The paper introducing GPT-2, demonstrating that scaling up a language model allows it to perform various downstream tasks zero-shot without explicit fine-tuning.
Read breakdownBERT
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Introduced bidirectional pretraining by masked language modelling, setting a new standard for natural language understanding tasks.
Read breakdownRoFormer
RoFormer: Enhanced Transformer with Rotary Position Embedding
Introduced Rotary Position Embedding (RoPE), a method that mathematically integrates absolute positional information with relative distances, becoming the standard for modern LLMs.
Read breakdownT5
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Introduced T5, a framework that reframes every NLP task into a text-to-text format, enabling a single model architecture to handle diverse tasks.
Read breakdownSwitch Transformer
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Scaled a sparse Mixture of Experts (MoE) model to a trillion parameters, proving that massive parameter scaling is possible without a proportional increase in compute.
Read breakdownMegatron-LM
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Introduced an elegant technique for Tensor Parallelism, allowing massive Transformer models to be split across multiple GPUs by partitioning the matrix multiplications themselves.
Read breakdownSequence to Sequence Learning
Sequence to Sequence Learning with Neural Networks
Introduced the Seq2Seq encoder-decoder architecture, allowing neural networks to map input sequences to output sequences of entirely different lengths.
Read breakdownDense Passage Retrieval (DPR)
Dense Passage Retrieval for Open-Domain Question Answering
The 2020 paper that proved dense neural embeddings could outperform traditional lexical search (like BM25) for open-domain question answering.
Read breakdownColBERT
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
The 2020 paper that introduced Late Interaction, bridging the speed of dual-encoders with the accuracy of cross-encoders for neural retrieval.
Read breakdownSentence-BERT (SBERT)
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
The 2019 paper that adapted BERT for generating semantically meaningful sentence embeddings that can be compared using cosine similarity.
Read breakdownMatryoshka Representation Learning
Matryoshka Representation Learning
The 2022 paper that introduced a way to train embedding models so their output vectors can be truncated to smaller sizes without retraining.
Read breakdownBillion-Scale Similarity Search with GPUs
Billion-Scale Similarity Search with GPUs
The 2017 paper by Facebook AI Research that introduced FAISS, making massive-scale vector similarity search practical.
Read breakdownHNSW
Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs
The 2016 paper that introduced the Navigable Small-World graph, the reigning standard for extremely fast approximate nearest neighbor search.
Read breakdownLatent Diffusion Models (LDM)
High-Resolution Image Synthesis with Latent Diffusion Models
The 2021 paper that brought diffusion models to the masses by running the generative process in a compressed latent space, drastically reducing compute requirements.
Read breakdownDDPM
Denoising Diffusion Probabilistic Models
The 2020 paper that proved diffusion models could generate high-quality images by learning to reverse a gradual noising process.
Read breakdownDDIM
Denoising Diffusion Implicit Models
The 2020 paper that dramatically accelerated diffusion model sampling, reducing generation time from thousands of steps to a few dozen without retraining.
Read breakdownDiffusion Transformers (DiT)
Scalable Diffusion Models with Transformers
The 2022 paper that proved Transformers could replace the U-Net as the backbone for diffusion models, unlocking predictable scaling laws for image generation.
Read breakdownConsistency Models
Consistency Models
The 2023 paper from OpenAI that enables high-quality diffusion generation in just a single step, bypassing the slow iterative sampling process entirely.
Read breakdownFlow Matching
Flow Matching for Generative Modeling
The 2022 paper that generalized diffusion models into a simpler, simulation-free framework based on Continuous Normalizing Flows and vector fields.
Read breakdownScore-Based SDEs
Score-Based Generative Modeling through Stochastic Differential Equations
The 2020 paper that unified diffusion models and score-matching models into a single, elegant framework based on continuous-time Stochastic Differential Equations (SDEs).
Read breakdownClassifier-Free Guidance (CFG)
Classifier-Free Diffusion Guidance
The 2022 paper that unlocked high-quality text-to-image generation by teaching diffusion models to heavily prioritize the text prompt over the unconditional image prior.
Read breakdownControlNet
Adding Conditional Control to Text-to-Image Diffusion Models
The 2023 paper that introduced a way to add precise spatial control (like edge maps or human poses) to large text-to-image diffusion models without retraining them.
Read breakdownVision Transformer (ViT)
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
The 2020 paper that proved Transformer architectures could replace CNNs for computer vision by treating image patches as a sequence of words.
Read breakdownSwin Transformer
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
The 2021 paper that brought hierarchical structure and local windows to Vision Transformers, making them practical for high-resolution tasks like object detection and segmentation.
Read breakdownDINOv2
DINOv2: Learning Robust Visual Features without Supervision
The 2023 paper from Meta that produced state-of-the-art self-supervised visual features, matching supervised models without using any labels or text.
Read breakdownSigLIP
Sigmoid Loss for Language Image Pre-Training
The 2023 paper that replaced CLIP's softmax contrastive loss with a simpler sigmoid loss, allowing for massive scaling and better performance.
Read breakdownSegment Anything Model (SAM)
Segment Anything
The 2023 Meta paper that introduced a promptable foundation model for image segmentation, capable of zero-shot segmentation of any object.
Read breakdownNeRF
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
The 2020 paper that revolutionized 3D computer vision by proving a neural network could "memorize" a 3D scene and render photorealistic novel views.
Read breakdown3D Gaussian Splatting
3D Gaussian Splatting for Real-Time Radiance Field Rendering
The 2023 paper that dethroned NeRFs for novel view synthesis by replacing the neural network with millions of explicit 3D splats, achieving real-time 1080p rendering.
Read breakdownGenerative Adversarial Networks (GAN)
Generative Adversarial Nets
Introduced the GAN, an elegant framework where two neural networks—a generator and a discriminator—compete against each other to create hyper-realistic synthetic data.
Read breakdownFID (Fréchet Inception Distance)
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
The 2017 paper that introduced Fréchet Inception Distance (FID), solving the critical problem of how to mathematically measure the quality of AI-generated images.
Read breakdownVariational Autoencoders (VAE)
Auto-Encoding Variational Bayes
Introduced the Variational Autoencoder, a generative model that learns a continuous, structured latent space through the mathematical 'reparameterization trick'.
Read breakdownVQ-VAE
Neural Discrete Representation Learning
The 2017 DeepMind paper that introduced the Vector Quantized Variational Autoencoder, turning continuous images into sequences of discrete, learnable 'tokens'.
Read breakdownU-Net
U-Net: Convolutional Networks for Biomedical Image Segmentation
Introduced U-Net, an elegant, symmetric architecture for image segmentation that remains the foundational backbone for modern diffusion models.
Read breakdownDeep Residual Learning
Deep Residual Learning for Image Recognition
The 2015 paper that introduced residual skip connections, solving the vanishing gradient problem and making extremely deep networks trainable.
Read breakdownConvNeXt
A ConvNet for the 2020s
The 2022 paper that modernized the standard ResNet architecture to prove CNNs could still match or beat Vision Transformers on accuracy and scalability.
Read breakdownEfficientNet
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
The 2019 paper that introduced compound scaling, a principled way to scale up convolutional networks across depth, width, and resolution.
Read breakdownImageNet Classification with Deep CNNs
ImageNet Classification with Deep Convolutional Neural Networks
The 2012 paper (AlexNet) that proved deep convolutional neural networks trained on GPUs could crush traditional computer vision methods.
Read breakdownCLIP
Learning Transferable Visual Models From Natural Language Supervision
The 2021 OpenAI paper that aligned text and images in a shared embedding space, unlocking zero-shot classification and the modern generative image era.
Read breakdownWhisper
Robust Speech Recognition via Weak Supervision
The 2022 OpenAI paper that achieved human-level robustness in speech recognition by training on a massive 680,000-hour dataset of noisy, weakly supervised web audio.
Read breakdownwav2vec 2.0
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
The 2020 Facebook AI paper that successfully applied self-supervised learning to audio, allowing speech recognition models to be trained with vastly less human-transcribed data.
Read breakdownAlphaFold 2
Highly Accurate Protein Structure Prediction with AlphaFold
The 2021 DeepMind paper that solved a 50-year-old grand challenge in biology by using AI to predict the 3D structure of a protein from its 1D amino acid sequence.
Read breakdownProximal Policy Optimization (PPO)
Proximal Policy Optimization Algorithms
Introduced PPO, an incredibly stable and sample-efficient Reinforcement Learning algorithm that became the default standard, eventually powering RLHF in ChatGPT.
Read breakdownDeep Double Descent
Deep Double Descent: Where Bigger Models and More Data Hurt
The 2019 paper that broke classical statistical theory, proving that making a neural network "too big" actually causes its error rate to drop again after an initial spike.
Read breakdownLottery Ticket Hypothesis
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
The 2018 MIT paper that proved massive neural networks contain tiny, sparse subnetworks that can learn just as fast and achieve the exact same accuracy as the giant model.
Read breakdownScaling Monosemanticity
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Anthropic's 2024 paper successfully applied Sparse Autoencoders to Claude 3 Sonnet, extracting millions of high-level, interpretable concepts (like 'Golden Gate Bridge' or 'security vulnerabilities') from a state-of-the-art production model.
Read breakdownTowards Monosemanticity
Towards Monosemanticity: Decomposing Neural Networks with Sparse Autoencoders
A 2023 mechanistic interpretability paper by Anthropic that successfully used Sparse Autoencoders to extract understandable, human-readable concepts from the impenetrable 'black box' of neural network activations.
Read breakdownSHAP (SHapley Additive exPlanations)
A Unified Approach to Interpreting Model Predictions
Introduced SHAP, a unified framework for interpreting complex machine learning models by applying game theory to calculate the exact marginal contribution of each feature.
Read breakdownLIME
Why Should I Trust You?: Explaining the Predictions of Any Classifier
The 2016 paper that introduced Local Interpretable Model-agnostic Explanations, a method to figure out exactly why a "black box" AI made a specific decision.
Read breakdownAdversarial Examples (FGSM)
Explaining and Harnessing Adversarial Examples
The 2014 paper by Goodfellow et al. that exposed a terrifying flaw in neural networks: adding invisible, mathematically calculated noise to an image causes the model to confidently misclassify it.
Read breakdownData Extraction Attacks
Extracting Training Data from Large Language Models
The 2020 paper that exposed a critical privacy flaw in LLMs, proving that massive generative models verbatim memorize their training data, which can be easily extracted by users.
Read breakdownConcrete Problems in AI Safety
Concrete Problems in AI Safety
The 2016 paper that defined the modern field of AI Alignment by clearly categorizing how optimization processes can go disastrously wrong in the real world.
Read breakdownDP-SGD
Deep Learning with Differential Privacy
The 2016 Google paper that proved you can train deep neural networks while mathematically guaranteeing the privacy of the individuals in the training dataset.
Read breakdownAdam Optimizer
Adam: A Method for Stochastic Optimization
Introduced Adam, an adaptive optimization algorithm that combined the best properties of momentum and RMSProp, becoming the default optimizer for deep learning.
Read breakdownBatch Normalization
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Introduced Batch Normalization, a technique that normalized activations across a mini-batch, making networks faster and dramatically more stable to train.
Read breakdownLayer Normalization
Layer Normalization
Introduced Layer Normalization, which normalizes activations across the features of a single data point rather than across a batch, making it ideal for sequence models.
Read breakdownGQA
GQA: Training Generalized Multi-Query Attention
Introduced Grouped-Query Attention, striking an optimal balance between the high quality of Multi-Head Attention and the fast inference speed of Multi-Query Attention.
Read breakdownDropout
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Introduced Dropout, a remarkably simple yet highly effective regularization technique that prevents neural networks from overfitting by randomly disabling neurons.
Read breakdownKnowledge Distillation
Distilling the Knowledge in a Neural Network
Formalized Knowledge Distillation, a method for training small, fast 'student' models to mimic the nuanced behavior of massive 'teacher' models.
Read breakdownWord2Vec
Efficient Estimation of Word Representations in Vector Space
Introduced Word2Vec, a highly efficient method for learning dense word embeddings that captured semantic meaning and analogical relationships.
Read breakdownRandom Forests
Random Forests
Formalized Random Forests, an ensemble learning method that combines bagging and random feature selection to build highly robust and accurate predictive models.
Read breakdownXGBoost
XGBoost: A Scalable Tree Boosting System
Introduced XGBoost, a highly scalable and regularized gradient boosting library that utterly dominated Kaggle competitions and tabular data modeling.
Read breakdownConformal Prediction
A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification
The 2021 tutorial that popularized Conformal Prediction, a mathematical framework to add rigorous uncertainty bounds (confidence intervals) to any machine learning model.
Read breakdownDouble/Debiased ML
Double/Debiased Machine Learning for Treatment and Structural Parameters
The 2016 economics and statistics paper that proved how to use highly flexible machine learning models to infer true causal effects, without bias.
Read breakdownALiBi: Attention with Linear Biases
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
ALiBi replaces standard positional embeddings with a simple linear penalty on attention scores, allowing language models to extrapolate to sequences much longer than they were trained on.
Read breakdownPAC Learning
A Theory of the Learnable
The 1984 paper that mathematically defined what it actually means for a computer to 'learn', laying the foundational theory for modern machine learning.
Read breakdownSpider
Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task
The 2018 dataset paper that defined the cross-domain text-to-SQL task and became the standard benchmark for natural language database querying.
Read breakdownLLaVA
Visual Instruction Tuning
The 2023 paper that demonstrated how to turn an open-source text LLM into a powerful multimodal model by projecting visual features into the text token space.
Read breakdownModel Cards
Model Cards for Model Reporting
The 2018 paper that established the industry standard for AI transparency, proposing that every machine learning model should be accompanied by a 'nutrition label' detailing its performance, limits, and biases.
Read breakdownZeRO
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Introduced the Zero Redundancy Optimizer, a memory optimization technology that partitions model states across GPUs, making it possible to train models with hundreds of billions of parameters.
Read breakdownParameter-Efficient Transfer Learning for NLP
Parameter-Efficient Transfer Learning for NLP
Introduces adapter modules, a parameter-efficient alternative to full fine-tuning that adds only a few trainable parameters per task while keeping the pre-trained model weights frozen.
Read breakdownMastering the Game of Go without Human Knowledge (AlphaZero)
Mastering the game of Go without human knowledge
Introduced a reinforcement learning algorithm that achieved superhuman performance in Go entirely through self-play, without any human data.
Read breakdownBLIP-2: Bootstrapping Language-Image Pre-training
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Proposes a compute-efficient pre-training strategy for vision-language models that bootstraps from off-the-shelf frozen image encoders and frozen large language models using a lightweight Querying Transformer.
Read breakdownHierarchical Text-Conditional Image Generation with CLIP Latents
Hierarchical Text-Conditional Image Generation with CLIP Latents
DALL-E 2 produces high-resolution, realistic images from text prompts by combining CLIP's text-image representations with a two-stage diffusion process.
Read breakdownDensely Connected Convolutional Networks
Densely Connected Convolutional Networks
Introduces DenseNets, where each layer is directly connected to every other layer in a feed-forward fashion, improving feature reuse and mitigating the vanishing gradient problem.
Read breakdownEnd-to-End Object Detection with Transformers
End-to-End Object Detection with Transformers
Treated object detection as a direct set prediction problem, using a Transformer encoder-decoder architecture to eliminate hand-designed anchors and non-maximum suppression.
Read breakdownPlaying Atari with Deep Reinforcement Learning (DQN)
Playing Atari with Deep Reinforcement Learning
Introduced Deep Q-Networks (DQN), proving that a single neural network architecture can learn to play multiple Atari games directly from raw pixel inputs.
Read breakdownSemi-Supervised Classification with Graph Convolutional Networks
Semi-Supervised Classification with Graph Convolutional Networks
Introduces a scalable approach for semi-supervised learning on graph-structured data based on an efficient variant of convolutional neural networks that operate directly on graphs.
Read breakdownGoing Deeper with Convolutions
Going Deeper with Convolutions
GoogLeNet introduced the Inception module, enabling networks to grow significantly deeper and wider while maintaining a strictly constrained computational budget.
Read breakdownImproving Language Understanding by Generative Pre-Training
Improving Language Understanding by Generative Pre-Training
Showed that training a Transformer language model on unlabeled text, followed by task-specific fine-tuning, dramatically improves performance across NLP tasks.
Read breakdownLong Short-Term Memory
Long Short-Term Memory
Introduced the LSTM architecture to solve the vanishing gradient problem in recurrent neural networks, enabling the learning of long-term dependencies.
Read breakdownMask R-CNN
Mask R-CNN
A conceptually simple, flexible, and general framework for object instance segmentation that dominated computer vision for years.
Read breakdownMasked Autoencoders Are Scalable Vision Learners
Masked Autoencoders Are Scalable Vision Learners
MAE proved that masking random patches of an image and reconstructing them is a highly scalable and effective method for self-supervised computer vision.
Read breakdownFast Transformer Decoding: One Write-Head is All You Need
Fast Transformer Decoding: One Write-Head is All You Need
Introduced Multi-Query Attention, drastically reducing memory bandwidth requirements during Transformer decoding by sharing key and value heads.
Read breakdownPrefix-Tuning: Optimizing Continuous Prompts for Generation
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Proposes a lightweight alternative to fine-tuning for natural language generation tasks, which keeps language model parameters frozen, but optimizes a small continuous task-specific vector (prefix).
Read breakdownRich Feature Hierarchies for Accurate Object Detection
Rich feature hierarchies for accurate object detection and semantic segmentation
Replaced complex ensemble detection systems with a simple combination of region proposals and Convolutional Neural Networks, revolutionizing object detection.
Read breakdownEfficiently Modeling Long Sequences with Structured State Spaces (S4)
Efficiently Modeling Long Sequences with Structured State Spaces
Introduced the Structured State Space (S4) model, enabling the efficient processing of extremely long sequences without the quadratic cost of attention.
Read breakdownOutrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Sparse Mixture of Experts allows neural networks to scale parameters enormously without proportional computational cost by routing inputs to specialized sub-networks.
Read breakdownA Style-Based Generator Architecture for GANs
A Style-Based Generator Architecture for GANs
StyleGAN revolutionized image generation by disentangling latent spaces and enabling fine-grained control over generated features at different scales.
Read breakdownSupport Vector Networks (Cortes & Vapnik)
Support-Vector Networks
Formalized the Support Vector Machine (SVM) algorithm for non-linear classification, defining one of the most powerful algorithms of the pre-deep learning era.
Read breakdownVisualizing Data using t-SNE
Visualizing Data using t-SNE
Introduced t-SNE, a highly effective dimensionality reduction technique that revolutionized the visualization of high-dimensional datasets.
Read breakdownToy Models of Superposition
Toy Models of Superposition
Anthropic's exploration of how neural networks can represent more features than they have dimensions by packing them into a superposition state.
Read breakdownVery Deep Convolutional Networks for Large-Scale Image Recognition
Very Deep Convolutional Networks for Large-Scale Image Recognition
VGG proved that increasing depth using homogeneous 3x3 convolutional filters was crucial for improving accuracy in image recognition.
Read breakdownWasserstein GAN
Wasserstein GAN
Introduced a new loss function for Generative Adversarial Networks based on the Earth Mover's distance, dramatically improving training stability.
Read breakdownYou Only Look Once
You Only Look Once: Unified, Real-Time Object Detection
Reframed object detection as a single regression problem, allowing a single neural network to predict bounding boxes and classes in real-time.
Read breakdownDecoupled Weight Decay Regularization (AdamW)
Decoupled Weight Decay Regularization
AdamW fixes a fundamental flaw in how Adam applies weight decay, proving that L2 regularization and weight decay are not equivalent for adaptive optimizers.
Read breakdownGloVe
GloVe: Global Vectors for Word Representation
Introduced GloVe, an unsupervised algorithm that learns dense word vectors by factoring a global word-word co-occurrence matrix, combining matrix factorization and local context windows.
Read breakdownNeural Machine Translation by Jointly Learning to Align and Translate
Neural Machine Translation by Jointly Learning to Align and Translate
Introduces the attention mechanism to sequence-to-sequence models, allowing networks to dynamically align inputs and outputs, removing the single vector bottleneck.
Read breakdownRethinking Positional Encoding (RoPE)
RoFormer: Enhanced Transformer with Rotary Position Embedding
An exploration of Rotary Position Embedding (RoPE), detailing how it mathematically unifies absolute and relative positional encoding for Transformers.
Read breakdownLongformer & Big Bird
Longformer: The Long-Document Transformer
Replaces quadratic self-attention with sparse local, global, and random connections to process sequences of up to 16,384 tokens.
Read breakdownReformer
Reformer: The Efficient Transformer
Reformer reduces Transformer complexity from O(N²) to O(N log N) using Locality-Sensitive Hashing for attention and Reversible Layers for memory efficiency.
Read breakdownRWKV
RWKV: Reinventing RNNs for the Transformer Era
Introduced a linear attention architecture that trains in parallel like a Transformer but performs inference as an efficient RNN with a fixed state.
Read breakdownRetNet
Retentive Network: A Successor to Transformer for Large Language Models
Introduced a novel architecture that replaces softmax attention with a retention mechanism, enabling parallel training and O(1) inference cost.
Read breakdownGShard
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
GShard introduces conditional computation and automatic sharding, enabling the training of a 600-billion parameter Transformer using sparsely-gated MoEs.
Read breakdownGroup Relative Policy Optimization (GRPO)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Introduced GRPO, a memory-efficient reinforcement learning algorithm that eliminates the need for a Critic model by computing relative advantages from a group of outputs.
Read breakdownSelf-Rewarding Language Models
Self-Rewarding Language Models
An approach where LLMs act as their own reward models, enabling iterative self-improvement without human-bottlenecked preference data.
Read breakdownModel-Agnostic Meta-Learning (MAML)
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
Introduced MAML, a meta-learning algorithm that explicitly trains a model's initial parameters so that a few gradient steps on a new task produce maximal performance.
Read breakdownIdentity Preference Optimization (IPO)
A General Theoretical Paradigm to Understand Learning from Human Preferences
A general theoretical paradigm that frames LLM alignment as game-theoretic preference optimization, fixing DPO's overfitting issues with a simpler regularized objective.
Read breakdownPerceiver and Perceiver IO
Perceiver IO: A General Architecture for Structured Inputs & Outputs
A general architecture that decouples input size from compute complexity by cross-attending high-dimensional byte arrays into a fixed, small set of latents.
Read breakdownImage-to-Image Translation with Conditional Adversarial Networks
Image-to-Image Translation with Conditional Adversarial Networks
Pix2Pix introduced a general-purpose framework for image-to-image translation using conditional GANs, replacing task-specific hand-engineered loss functions.
Read breakdownUnpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks
Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks
Introduced cycle consistency loss to enable image-to-image translation without needing paired examples in the training dataset.
Read breakdownProgressive Growing of GANs for Improved Quality, Stability, and Variation
Progressive Growing of GANs for Improved Quality, Stability, and Variation
A training method for GANs that starts with low-resolution images and progressively adds layers to increase resolution, improving stability and quality.
Read breakdownStyleGAN2
Analyzing and Improving the Image Quality of StyleGAN
StyleGAN2 eliminates droplet artifacts by replacing AdaIN with weight demodulation and improves latent interpolation with path length regularization.
Read breakdownBigGAN
Large Scale GAN Training for High Fidelity Natural Image Synthesis
Scaled up Generative Adversarial Networks to create high-resolution, high-fidelity images using large batch sizes and the truncation trick.
Read breakdownTaming Transformers for High-Resolution Image Synthesis
Taming Transformers for High-Resolution Image Synthesis
Introduces VQGAN, combining the local efficiency of convolutional models with the global expressivity of Transformers through a learned discrete codebook.
Read breakdownGLIDE
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
A generative diffusion model combining classifier-free guidance with text embeddings, establishing the foundation for DALL-E 2's photorealism.
Read breakdownImagen
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
Imagen combines large frozen text encoders like T5 with cascaded diffusion models for state-of-the-art photorealistic text-to-image generation.
Read breakdownStable Diffusion 3
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
The 2024 architecture introducing Multimodal Diffusion Transformers (MM-DiT) and Rectified Flow matching to vastly improve prompt adherence and text generation.
Read breakdownDreamBooth
DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
A method to personalize text-to-image diffusion models using just 3-5 images, binding a specific subject to a rare token while preserving class prior.
Read breakdownTextual Inversion
An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
Personalizing text-to-image models by learning a new pseudo-word token embedding to capture a specific subject, keeping the base model completely frozen.
Read breakdownInstructPix2Pix
InstructPix2Pix: Learning to Follow Image Editing Instructions
An instruction-following image editing model trained on a synthetic dataset generated by combining GPT-3 and Stable Diffusion.
Read breakdownPrompt-to-Prompt Image Editing with Cross Attention Control
Prompt-to-Prompt Image Editing with Cross Attention Control
Edit images in text-to-image diffusion models without fine-tuning by injecting and manipulating cross-attention maps during the generation process.
Read breakdownNeural Style Transfer
A Neural Algorithm of Artistic Style
The seminal paper that separated image content from artistic style using deep convolutional networks, enabling automated style transfer without spatial constraints.
Read breakdownResNeXt
Aggregated Residual Transformations for Deep Neural Networks
Introduced cardinality as a new dimension for scaling neural networks, splitting convolutions into parallel grouped pathways.
Read breakdownModernBERT
ModernBERT: Dragging BERT into the 2020s
Updating the classic encoder architecture with RoPE, GLU, and FlashAttention to make it context-aware up to 8k tokens and extremely fast.
Read breakdownReinforcement Learning with Calibrated Decisions (RLCD)
RLCD: Reinforcement Learning with Calibrated Decisions for System 1 Routers
Training non-generative decision models to output honest probabilities by using proper scoring rules as the reinforcement learning reward signal.
Read breakdownMobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
Introduced depthwise separable convolutions to drastically reduce the computational cost and model size of vision networks, enabling on-device ML.
Read breakdownEfficientNetV2: Smaller Models and Faster Training
EfficientNetV2: Smaller Models and Faster Training
Introduced Fused-MBConv and progressive learning to make the EfficientNet family train much faster and run with fewer parameters.
Read breakdownGrounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Fuses vision and language representations by injecting text prompt features deep into the object detector, enabling open-set detection of anything described in text.
Read breakdownOWL-ViT: Vision Transformer for Open-World Localization
Simple Open-Vocabulary Object Detection with Vision Transformers
Adapts standard Vision Transformers into zero-shot object detectors by attaching light prediction heads and scaling pre-training on image-text pairs.
Read breakdownImageBind: One Embedding Space To Bind Them All
ImageBind: One Embedding Space To Bind Them All
Learns a single joint embedding space for six different modalities using image-paired data, enabling emergent cross-modal alignment without explicit pairing.
Read breakdownVideo Diffusion Models
Video Diffusion Models
Extended image diffusion architectures into the time domain using 3D U-Nets, enabling the first wave of high-fidelity generative video.
Read breakdownVideoPoet: A Large Language Model for Zero-Shot Video Generation
VideoPoet: A Large Language Model for Zero-Shot Video Generation
Pioneered modeling video generation as a pure language modeling task using a standard Transformer, tokenizing video and audio into a unified sequence.
Read breakdownMusicGen: Simple and Controllable Music Generation
Simple and Controllable Music Generation
A single-stage autoregressive transformer that generates high-fidelity music from text prompts using an efficient interleaved acoustic tokenization scheme.
Read breakdownWaveNet: A Generative Model for Raw Audio
WaveNet: A Generative Model for Raw Audio
Pioneered the generation of high-fidelity speech and audio by autoregressively predicting raw audio waveforms sample-by-sample using dilated causal convolutions.
Read breakdownQwen2.5: Foundation Models and Vision-Language Extensions
Qwen2.5 Technical Report
A suite of highly capable open-weight models achieving state-of-the-art performance across reasoning, coding, and vision tasks through extensive pre-training and alignment.
Read breakdownGemini 2.5 Technical Report: Spatial Agents and Multimodal Advances
Gemini 2.5: Advanced Capabilities in Spatial and Multimodal Reasoning
Introduced advanced spatial understanding and agentic workflows to the Gemini architecture, heavily optimizing for complex coding and multimodal reasoning.
Read breakdownDeepSeek-V3 Technical Report
DeepSeek-V3 Technical Report
An incredibly efficient Mixture-of-Experts architecture that pushed the limits of training cost-efficiency while matching state-of-the-art reasoning benchmarks.
Read breakdownGLM-4 Technical Report
GLM-4 Technical Report
A robust bilingual foundation model emphasizing agentic tool use, massive context processing, and multimodal integration out of the box.
Read breakdownFlamingo: a Visual Language Model for Few-Shot Learning
Flamingo: a Visual Language Model for Few-Shot Learning
Pioneered few-shot multimodal learning by injecting visual tokens directly into a frozen LLM using gated cross-attention, allowing it to adapt to new tasks rapidly.
Read breakdown