Natural Language Processing
30 interview questions in this topic, each with its full answer shown below. Use "Collapse all" to skim just the titles.
“What is Bag-of-Words, and what are its drawbacks?”
Bag-of-Words (BoW) represents text as a frequency count of words from a vocabulary, ignoring grammar and order. Its main drawbacks are loss of semantic meaning, inability to capture context, and creating high-dimensional, sparse vectors.
Answer
The Bag-of-Words (BoW) model is one of the simplest methods for feature extraction in Natural Language Processing. It converts text into numerical vectors that can be used by machine learning algorithms.
How it works:
- Vocabulary Creation: It scans the entire training corpus and creates a vocabulary of all unique words.
- Vectorization: For each document (or sentence), it creates a vector of the same length as the vocabulary. Each element in the vector represents the frequency (count) of the corresponding word in that specific document.
It is called a "bag" of words because any information about the order or structure of words in the document is discarded. The model only cares about whether known words occur in the document, not where they occur in the document.
Drawbacks of Bag-of-Words:
- Loss of Semantic Meaning and Context: Because word order is ignored, "The dog bit the man" and "The man bit the dog" result in the exact same vector representation, despite having drastically different meanings. It cannot capture phrases or contextual nuances.
- Sparsity: In a large corpus, the vocabulary can grow to tens or hundreds of thousands of words. However, a single short document will only contain a tiny fraction of these words. This results in highly sparse vectors (mostly zeroes), which are computationally expensive and inefficient for many models.
- Out of Vocabulary (OOV) Issues: If the model encounters a word during inference that was not in the training vocabulary, it simply ignores it, potentially losing valuable information.
- Equal Weighting: It treats all words equally based on frequency. Frequent but uninformative words can dominate the representation, which is why BoW is often paired with stop-word removal or upgraded to TF-IDF.
💡 Note While BoW is outdated for complex tasks, it remains a surprisingly strong baseline for simple text classification tasks like spam detection, primarily due to its simplicity and speed.
“Compare BERT, GPT, and T5 pretraining objectives and the tasks each suits best.”
BERT is an encoder pre-trained via Masked Language Modeling, ideal for NLU tasks. GPT is a decoder pre-trained via Causal Language Modeling, ideal for generation. T5 is an encoder-decoder pre-trained by masking spans, converting all NLP problems into text-to-text tasks.
Answer
The Transformer architecture spawned three primary paradigms for pre-training large language models, each tailored for different downstream tasks.
1. BERT (Bidirectional Encoder Representations from Transformers)
- Architecture: Encoder-only.
- Pre-training Objective: Masked Language Modeling (MLM). It randomly masks 15% of the input tokens and trains the model to predict them based on the bidirectional context (words to the left and right). It also uses Next Sentence Prediction (NSP).
- Best Suited For: Natural Language Understanding (NLU) tasks that require deep comprehension of the whole sequence, such as text classification, sentiment analysis, NER, and extractive question answering.
2. GPT (Generative Pre-trained Transformer)
- Architecture: Decoder-only.
- Pre-training Objective: Causal (or Autoregressive) Language Modeling. It predicts the next word in a sequence given all previous words. It only has access to the leftward context (unidirectional).
- Best Suited For: Natural Language Generation (NLG) tasks. Because it is fundamentally designed to predict the next word, it excels at open-ended text generation, conversational AI, and zero/few-shot prompting.
3. T5 (Text-to-Text Transfer Transformer)
- Architecture: Encoder-Decoder (Seq2Seq).
- Pre-training Objective: Corrupted Span Replacement. It randomly masks contiguous spans of text and trains the decoder to output just the masked spans.
- Best Suited For: The core philosophy of T5 is that every NLP task can be framed as a text-to-text problem. It excels at translation, abstractive summarization, and complex generation tasks, but can also perform classification by generating the word "positive" or "negative".
💡 Note While BERT requires adding task-specific classification heads for fine-tuning, T5 and modern GPT variants handle multiple tasks purely through prompting or prompt-tuning without changing the architecture.
“Explain the difference between BLEU, ROUGE, and METEOR, and what each is best used for.”
BLEU measures precision-based n-gram overlap (good for translation). ROUGE measures recall-based n-gram overlap (good for summarization). METEOR improves on BLEU by aligning words using synonyms and stemming, offering better correlation with human judgment.
Answer
Evaluating generative NLP tasks (like translation or summarization) is notoriously difficult because there are many valid ways to express the same idea. BLEU, ROUGE, and METEOR are automated metrics designed to compare model-generated text against human reference texts.
1. BLEU (Bilingual Evaluation Understudy)
- Focus: Precision-based. It asks, "How many n-grams in the generated text appear in the reference text?"
- Mechanism: It calculates modified n-gram precision and applies a brevity penalty to prevent the model from gaming the system by outputting very short, highly precise translations.
- Best Used For: Machine Translation. It is the oldest and most standard metric, though it struggles with synonyms and phrasing variations.
2. ROUGE (Recall-Oriented Understudy for Gisting Evaluation)
- Focus: Recall-based. It asks, "How many n-grams in the reference text were successfully captured by the generated text?"
- Mechanism: ROUGE-N measures n-gram overlap, while ROUGE-L measures the Longest Common Subsequence (preserving word order).
- Best Used For: Text Summarization. In summarization, it is crucial to capture all the important information from the source (recall), making ROUGE the standard metric.
3. METEOR (Metric for Evaluation of Translation with Explicit ORdering)
- Focus: Harmonic mean of precision and recall (weighted toward recall), with explicit semantic matching.
- Mechanism: Unlike BLEU which requires exact string matches, METEOR aligns words using exact matches, stemming (matching "run" to "running"), and synonyms (using WordNet). It also includes a fragmentation penalty for bad word order.
- Best Used For: Machine Translation and general generation tasks where correlation with human judgment is prized over pure exact-match string metrics.
💡 Note All three metrics suffer from being lexical (surface-level) metrics. They cannot truly judge meaning. Modern evaluation is moving toward model-based metrics like BERTScore, which use contextual embeddings to measure semantic similarity.
“How does Byte-Pair Encoding (BPE) work, and why is subword tokenization preferred?”
BPE starts with a character-level vocabulary and iteratively merges the most frequent adjacent character pairs into new subword tokens. It is preferred because it balances vocabulary size, captures semantic subcomponents, and elegantly handles out-of-vocabulary words.
Answer
Byte-Pair Encoding (BPE) is a data compression algorithm that has been adapted into the standard subword tokenization technique for modern Large Language Models (like GPT and RoBERTa).
How BPE Works:
- Initialization: The algorithm starts by splitting all words in the training corpus into individual characters (or bytes) and adds a special end-of-word symbol. The initial vocabulary is just the set of all unique characters.
- Frequency Counting: It counts the frequency of all adjacent token pairs in the corpus.
- Merging: It identifies the most frequently occurring pair of adjacent tokens and merges them into a single, new token. This new token is added to the vocabulary.
- Iteration: Steps 2 and 3 are repeated for a pre-defined number of merge operations.
For example, if the corpus frequently contains "l o w", the pair "l" and "o" might be merged into "lo", and eventually "lo" and "w" into the subword "low".
Why Subword Tokenization is Preferred:
- Solves Out-Of-Vocabulary (OOV) Issues: Word-level tokenization fails when it encounters a word not in its training data (resulting in an
<UNK>token). BPE guarantees it can represent any text because, in the worst case, it can fall back to the base character tokens to construct the unknown word. - Vocabulary Efficiency: It keeps the vocabulary size manageable. Instead of storing every conjugation of a word, it can store base stems and suffixes separately (e.g., "play", "##ing", "##ed").
- Captures Morphology: It naturally breaks morphologically complex words into meaningful subwords, allowing the model to infer meaning for novel compound words.
💡 Note WordPiece (used by BERT) is very similar to BPE. The main difference is that while BPE merges the most frequent pairs, WordPiece merges the pair that maximizes the likelihood of the language model (highest mutual information).
“What are the main steps of a classic NLP pipeline?”
A classic NLP pipeline consists of text acquisition, preprocessing (cleaning, tokenization, stopping, stemming/lemmatization), feature engineering (BoW, TF-IDF), modeling (ML algorithms), and evaluation.
Answer
A classic Natural Language Processing (NLP) pipeline typically involves several distinct stages to transform raw text into a format suitable for machine learning models and to extract meaningful insights.
The main steps are:
- Text Acquisition: Gathering the raw textual data from sources like databases, web scraping, or APIs.
- Text Cleaning: Removing unnecessary elements like HTML tags, special characters, URLs, and standardizing casing (usually converting to lowercase).
- Preprocessing: This involves several sub-steps to break down and normalize the text:
- Tokenization: Splitting the text into individual words or subwords (tokens).
- Stop Word Removal: Filtering out common but uninformative words (e.g., "the", "is", "a").
- Stemming/Lemmatization: Reducing words to their root or base form (e.g., "running" to "run").
- Feature Engineering (Vectorization): Converting the preprocessed text into numerical representations that machine learning models can understand. Common techniques include Bag-of-Words (BoW) and Term Frequency-Inverse Document Frequency (TF-IDF).
- Modeling: Training a machine learning model (like Naive Bayes, SVM, or Logistic Regression) on the numerical features to perform a specific task, such as text classification or sentiment analysis.
- Evaluation and Deployment: Assessing the model's performance using metrics like accuracy, precision, recall, or F1-score, and finally deploying it for inference.
💡 Note Modern deep learning approaches often combine feature engineering and modeling into a single step using neural networks and word embeddings, skipping steps like explicit stemming or BoW.
“How do contextual embeddings (e.g., BERT) differ from static embeddings?”
Static embeddings assign one fixed vector to a word regardless of context. Contextual embeddings generate dynamic vectors for words at inference time based on the surrounding sentence, allowing them to differentiate word meanings (polysemy).
Answer
The evolution from static to contextual embeddings represents one of the most significant leaps in Natural Language Processing, solving the critical problem of polysemy (words with multiple meanings).
Static Embeddings (e.g., Word2Vec, GloVe, FastText): In static embedding models, each word in the vocabulary is assigned exactly one fixed, unchanging vector representation during training.
- The Limitation: Because the vector is fixed, the model cannot distinguish between different meanings of the same word. The word "bank" in "river bank" and "bank account" receives the exact same vector representation. The embedding is essentially an average of all the contexts the word appeared in during training.
Contextual Embeddings (e.g., ELMo, BERT, GPT): Contextual embeddings do not look up a fixed vector from a table. Instead, the embedding for a word is generated dynamically at inference time by passing the entire sentence through a deep neural network (usually a Transformer).
- The Advantage: The model calculates the representation of a word based on the words surrounding it. Therefore, the vector for "bank" in "river bank" will be fundamentally different from the vector for "bank" in "bank account."
- Mechanism: Models like BERT use the self-attention mechanism to weigh the importance of all other words in the sentence when constructing the embedding for the current word, ensuring deeply contextualized semantic meaning.
💡 Note While ELMo used bi-directional LSTMs to achieve contextual embeddings, modern architectures rely entirely on Transformers. Static embeddings are still useful for lightweight, low-resource tasks where deploying a heavy Transformer is unfeasible.
“Explain how coreference resolution works and why it is hard.”
Coreference resolution identifies all expressions in a text that refer to the same real-world entity. It is difficult because it requires world knowledge, context understanding, gender/number agreement, and dealing with ambiguous pronouns (like the Winograd Schema Challenge).
Answer
Coreference Resolution is the NLP task of finding all linguistic expressions (mentions) in a given text that refer to the same real-world entity. For example, in the sentence: "Barack Obama said he would sign the bill, but the former president later vetoed it." A coreference system must group former president into one cluster, and it into another.
How it Works: Modern systems typically treat this as a clustering problem:
- Mention Detection: Identify all potential entities and pronouns using NER and POS tagging.
- Pairwise Scoring: A neural network (often Transformer-based) evaluates every pair of mentions and calculates the probability that they refer to the same entity.
- Clustering: Graph algorithms or greedy linking group the highly probable pairs into distinct clusters.
Why is it so hard?
- Requires World Knowledge: "The trophy didn't fit into the brown suitcase because it was too large." What does "it" refer to? The trophy. Now change "large" to "small". Now "it" refers to the suitcase. This is known as a Winograd Schema, and it requires common sense reasoning, not just syntax parsing.
- Long-Range Dependencies: A pronoun in paragraph 3 might refer to a proper noun introduced in paragraph 1. Maintaining context over long documents is difficult.
- Cataphoric References: Usually, a pronoun refers back to an earlier noun (anaphora). Sometimes, it refers forward: "Even though he was tired, John kept running." Resolving this requires looking ahead.
- Implicit References: Sometimes the antecedent isn't explicitly stated but implied by an event, making standard mention-matching fail.
💡 Note Coreference resolution is a critical preprocessing step for Information Extraction and Question Answering. If a system doesn't know who "he" is, it cannot accurately answer questions about the text.
“How do Conditional Random Fields differ from HMMs for sequence labeling?”
HMMs are generative models that assume observation independence and model the joint probability of states and observations. CRFs are discriminative models that directly model the conditional probability of labels given the entire observation sequence, allowing for complex, overlapping features.
Answer
Hidden Markov Models (HMMs) and Conditional Random Fields (CRFs) are classical statistical methods for sequence labeling tasks like Part-of-Speech tagging and Named Entity Recognition (NER).
Hidden Markov Models (HMMs)
- Type: Generative model. It models the joint probability of the observation sequence (words, ) and the hidden state sequence (labels, ).
- Assumptions: HMMs rely on strict independence assumptions to make computations tractable. Specifically, it assumes that the current observation depends only on the current hidden state, ignoring the rest of the sentence.
- Limitation: Because of these strict assumptions, it is very difficult to include arbitrary, overlapping, or long-range features (like "is the next word capitalized?" or "does the previous word end in '-ing'?").
Conditional Random Fields (CRFs)
- Type: Discriminative model. It directly models the conditional probability of the label sequence given the entire observation sequence.
- Advantage: CRFs drop the strict independence assumptions of HMMs. By conditioning on the entire input sequence globally, CRFs can incorporate rich, complex, and highly overlapping features (e.g., character n-grams, capitalization, word shapes) from anywhere in the sentence to predict the current label.
- Performance: Because they optimize the conditional probability directly and allow for better feature engineering, CRFs historically vastly outperformed HMMs on NLP tasks.
💡 Note Even in the age of Deep Learning, CRFs remain highly relevant. Many state-of-the-art NER systems use an architecture like BiLSTM-CRF or BERT-CRF, where the neural network acts as a powerful feature extractor, and a CRF layer sits on top to ensure valid label transitions (e.g., ensuring an 'I-ORG' tag never follows a 'B-PER' tag).
“How would you detect and mitigate domain shift when a text classifier moves to a new domain?”
Detect domain shift by comparing feature distributions (KL divergence) or performance drops on a target validation set. Mitigate it using Domain Adaptation techniques like continued pre-training on target domain text, or adversarial training to learn domain-invariant features.
Answer
Domain shift occurs when a machine learning model is trained on data from one distribution (the source domain, e.g., formal news articles) but applied to data from a different distribution (the target domain, e.g., informal Twitter posts). This usually results in a severe drop in performance.
How to Detect Domain Shift:
- Performance Degradation: The most obvious sign is a significant drop in accuracy/F1-score when tested on a small labeled sample of the target domain.
- Data Distribution Metrics: Compare the vocabulary overlap, TF-IDF distributions, or embedding representations between the source and target data using metrics like Kullback-Leibler (KL) Divergence or Maximum Mean Discrepancy (MMD).
- Proxy Classifier: Train a simple classifier to predict whether a document came from the source or target domain. If it can do so with high accuracy, a domain shift exists.
How to Mitigate Domain Shift (Domain Adaptation):
- Continued Pre-training (Domain-Adaptive Pre-training - DAPT): If using a model like BERT, take the massive amount of unlabeled text from your target domain and run Masked Language Modeling on it. This teaches the model the specific vocabulary and syntax of the new domain before fine-tuning it on the source labels.
- Data Augmentation: Blend data. Manually label a small set of target data (Few-Shot learning) and mix it with the source data.
- Domain Adversarial Neural Networks (DANN): Train the model with a feature extractor and two heads: a classification head and a domain discriminator head. Use a Gradient Reversal Layer on the domain head to force the feature extractor to learn representations that are useful for classification but make it impossible to tell which domain the text came from.
💡 Note Modern Large Language Models exhibit strong zero-shot robustness to domain shift. Often, simply updating the prompt (e.g., "Act as a financial analyst and classify this tweet...") is sufficient to mitigate minor domain shifts without retraining.
“How would you design an entity-linking system that maps mentions to a knowledge graph?”
An entity linking system uses NER to find mentions, generates candidate entities from a knowledge graph (via string matching or dense retrieval), and then uses a cross-encoder model to disambiguate and select the correct entity based on the mention's context.
Answer
Entity Linking (also known as Named Entity Disambiguation) takes mentions extracted from text and maps them to unique identifiers in a Knowledge Graph (like Wikidata or DBpedia).
For example, mapping the mention "Apple" in "Apple reported record earnings" to the company (Q312), not the fruit (Q89).
Design of an Entity Linking System:
-
Mention Detection (NER): First, use a Named Entity Recognition (NER) model to identify spans of text that represent entities.
-
Candidate Generation: For each extracted mention, retrieve a subset of possible entities from the Knowledge Graph to reduce the search space from millions to a handful (e.g., top 100).
- Techniques: Use alias dictionaries, exact string matching, BM25 (TF-IDF search over entity descriptions), or Bi-Encoder dense retrieval (embedding the mention and doing a vector search against entity embeddings).
-
Candidate Disambiguation (Ranking): This is the core ML challenge. You must rank the generated candidates to pick the correct one.
- Techniques: Use a Cross-Encoder Transformer. Concatenate the sentence containing the mention with the description of the candidate entity from the Knowledge Graph:
[CLS] Text sentence [SEP] Candidate Description [SEP]. The model outputs a probability score of a match. - Features to leverage:
- Local Context: Does the sentence context match the entity description?
- Global Coherence: Do the other entities linked in the same document logically co-occur with this candidate? (e.g., if "Jobs" and "Cupertino" are also in the text, "Apple" the company is highly likely).
- Techniques: Use a Cross-Encoder Transformer. Concatenate the sentence containing the mention with the description of the candidate entity from the Knowledge Graph:
-
NIL Handling: If the highest-scoring candidate is below a certain threshold, the system must confidently predict "NIL" (Not In Lexicon), meaning the entity does not exist in the Knowledge Graph yet.
💡 Note BLINK, developed by Facebook Research, is a standard architecture for modern entity linking, utilizing a fast Bi-Encoder for candidate generation and a highly accurate Cross-Encoder for disambiguation.
“What is the difference between extractive and abstractive summarization?”
Extractive summarization selects and pieces together the most important existing sentences from the source text. Abstractive summarization generates new sentences, paraphrasing and condensing the core meaning like a human would.
Answer
Text summarization in NLP aims to condense a long document into a shorter version while preserving the core information. There are two fundamentally different approaches to achieving this: Extractive and Abstractive.
Extractive Summarization Think of this as using a highlighter. The system scores all the sentences in the document based on their importance (using algorithms like TextRank, TF-IDF, or modern deep learning classifiers). It then extracts the highest-scoring sentences verbatim and concatenates them to form the summary.
- Pros: Much easier to build, computationally cheaper, and guarantees grammatical correctness (since the sentences were written by the original author). It is very robust and won't hallucinate false information.
- Cons: Summaries can be disjointed, lack narrative flow, and cannot paraphrase or condense information if a key idea spans multiple long sentences.
Abstractive Summarization Think of this as reading a document, closing it, and explaining it to a friend in your own words. The system uses advanced sequence-to-sequence models (typically Large Language Models like BART, T5, or GPT) to comprehend the entire text and generate entirely new sentences that capture the essence of the document.
- Pros: Produces fluent, concise, and highly cohesive summaries. It can paraphrase and merge concepts effectively.
- Cons: Requires massive amounts of compute and training data. It is prone to "hallucinations"—generating facts that sound plausible but were not in the original text.
💡 Note Modern state-of-the-art summarization uses entirely abstractive methods powered by LLMs. However, for highly sensitive legal or medical documents where factual hallucination is unacceptable, extractive methods (or hybrid approaches) are still heavily utilized.
“How do you handle out-of-vocabulary words and misspellings in production text systems?”
Handle OOV words and misspellings by using subword tokenization (BPE/WordPiece), character-level embeddings (FastText), robust spelling correction pipelines (edit distance, SymSpell), and domain-specific vocabulary augmentation during training.
Answer
Out-of-Vocabulary (OOV) words and misspellings are ubiquitous in real-world, user-generated text (social media, search queries, customer support). If an NLP system only knows dictionary words, it will fail dramatically in production.
Here are the primary strategies to handle them:
1. Subword Tokenization (The Modern Standard)
Algorithms like Byte-Pair Encoding (BPE), WordPiece, or SentencePiece are the most robust defense against OOV words. If a model encounters a misspelled word like "amazzzing", it doesn't assign it an <UNK> (unknown) token. Instead, it breaks it down into known subwords (e.g., ["am", "##az", "##zz", "##ing"]). The model can often infer the meaning from these constituent parts.
2. Character-Level or Subword Embeddings If using static embeddings, use FastText instead of Word2Vec or GloVe. FastText represents words as bags of character n-grams. Therefore, it can generate a vector for a misspelled or novel word on the fly, which will naturally be geometrically close to the correctly spelled word.
3. Explicit Spelling Correction Pipelines Before feeding text to the model, run a fast spelling correction step:
- Edit Distance algorithms (Levenshtein): Find dictionary words closest to the misspelling.
- SymSpell: An extremely fast algorithm that uses symmetric delete spelling correction, ideal for low-latency production environments.
- Contextual Spell Checking: Using a lightweight language model to pick the correct spelling correction based on surrounding words.
4. Data Augmentation during Training Intentionally inject noise into your training data. Randomly apply character swaps, deletions, or insertions mimicking common typos to make the model robust to misspellings natively.
💡 Note Blindly correcting spelling can sometimes destroy information. In specialized domains (like gaming or crypto), intentional misspellings (e.g., "pwned", "HODL") have specific meanings that standard spell-checkers would ruin.
“How would you build a multilingual NLP system, and what problems arise with low-resource languages?”
Build a multilingual system using massively multilingual pre-trained models (like mBERT or XLM-R) which map multiple languages into a shared vector space. Low-resource languages suffer from poor tokenization, lack of training data, and token representation imbalance.
Answer
Building a multilingual NLP system today rarely involves training separate models from scratch for each language. Instead, the standard approach relies on zero-shot cross-lingual transfer using massively multilingual models.
How to Build It:
- Foundation: Start with a multilingual pre-trained Transformer like mBERT (Multilingual BERT) or XLM-RoBERTa. These models are pre-trained on corpora consisting of 100+ languages simultaneously without explicit translation pairs. They learn to map similar concepts from different languages into the same semantic embedding space.
- Fine-tuning: Fine-tune the model on your specific task (e.g., Sentiment Analysis) using your labeled data, which might only be in English (a high-resource language).
- Zero-Shot Transfer: Because the underlying representations are language-agnostic, the fine-tuned model can now infer sentiment in Spanish, Hindi, or Swahili with surprisingly high accuracy, despite never seeing labeled data for those languages during fine-tuning.
Problems with Low-Resource Languages: While this approach is powerful, it severely degrades for low-resource languages (languages with minimal digital footprint).
- The Curse of Multilinguality: A model has limited capacity. Accommodating 100 languages dilutes the performance on individual languages. High-resource languages dominate the training data, leading the model to underrepresent low-resource languages.
- Tokenization Disparity: Subword tokenizers (like BPE) trained on predominantly English corpora will break low-resource languages into heavily fragmented, meaningless characters, destroying semantic meaning.
- Syntactic Differences: Cross-lingual transfer struggles if the target low-resource language has a fundamentally different sentence structure (e.g., Subject-Object-Verb) than the high-resource training language.
💡 Note To mitigate low-resource issues, techniques like Cross-Lingual Alignment (using small bilingual dictionaries to align embedding spaces) and Adapter Modules (training tiny language-specific layers while freezing the main model) are highly effective.
“What are n-grams, and what is the role of smoothing in n-gram language models?”
An n-gram is a contiguous sequence of n items (usually words) from a text. Smoothing is critical in n-gram models to handle zero-frequency problems by reassigning some probability mass from seen n-grams to unseen n-grams, preventing the model from assigning zero probability to novel sentences.
Answer
What are n-grams? An n-gram is a contiguous sequence of n items (characters, syllables, or words) from a given sample of text or speech.
- Unigram (n=1): "The", "quick", "brown"
- Bigram (n=2): "The quick", "quick brown"
- Trigram (n=3): "The quick brown"
In traditional language modeling, n-grams are used to predict the next word in a sequence based on the Markov assumption: the probability of a word depends only on the previous words. For example, a bigram model calculates .
The Role of Smoothing: A major flaw with n-gram models is the zero-frequency problem. If a specific sequence of words never appeared in the training corpus, the model assigns it a probability of 0. Because sequence probabilities are calculated by multiplying the probabilities of their constituent n-grams, a single unseen n-gram will cause the probability of the entire sentence to become 0.
Smoothing solves this by taking a small amount of probability mass from frequently occurring n-grams and distributing it to n-grams with zero counts.
- Laplace (Add-One) Smoothing: The simplest form, adding 1 to the count of every possible n-gram.
- Kneser-Ney Smoothing: A more advanced and effective technique that considers the diversity of contexts a word appears in (continuation probability) rather than just its raw frequency.
💡 Note While modern neural language models (like Transformers) have largely replaced n-gram models for complex tasks, n-grams remain highly useful for fast, baseline text classification, spelling correction, and measuring evaluation metrics like BLEU.
“What is Named Entity Recognition (NER)?”
Named Entity Recognition (NER) is an information extraction task that seeks to locate and classify named entities mentioned in unstructured text into predefined categories such as person names, organizations, locations, medical codes, time expressions, and quantities.
Answer
Named Entity Recognition (NER) is a core task in Natural Language Processing focused on information extraction. It involves parsing through unstructured text to locate specific entities and categorize them into predefined classes.
Common predefined categories include:
- Person (PER): "Barack Obama", "Elon Musk"
- Organization (ORG): "Google", "United Nations"
- Location (LOC): "New York", "Mount Everest"
- Date/Time (DATE): "July 4th", "tomorrow"
- Miscellaneous (MISC): Currencies, percentages, nationalities.
Why is NER useful? NER transforms raw, unstructured text into structured data. This structured data is highly valuable for downstream applications:
- Information Retrieval: Enhancing search engines by allowing searches based on entities rather than just keyword matching.
- Content Recommendation: Tagging news articles with relevant entities (e.g., companies mentioned) to recommend similar content.
- Customer Support: Automatically extracting product names or issue types from customer tickets to route them to the correct department.
How it works: NER is typically framed as a sequence labeling task. The model must process a sequence of tokens and assign a label to each token. It usually employs the BIO tagging scheme (Begin, Inside, Outside) to handle entities that span multiple words (e.g., "B-ORG" for the first word of an organization, "I-ORG" for subsequent words).
Historically, algorithms like Conditional Random Fields (CRFs) were heavily used. Today, fine-tuning pre-trained language models like BERT on NER datasets provides state-of-the-art performance due to their deep contextual understanding.
💡 Note NER systems can be highly domain-specific. A model trained on news articles will likely perform poorly at extracting protein structures or gene names in biomedical literature. Domain adaptation or specialized models (like BioBERT) are required.
“What is Part-of-Speech tagging?”
Part-of-Speech (POS) tagging is the process of marking up a word in a text corpus as corresponding to a particular part of speech, such as noun, verb, adjective, or adverb, based on its definition and context.
Answer
Part-of-Speech (POS) tagging is a fundamental Natural Language Processing task that involves assigning a grammatical category (like noun, verb, adjective, adverb, pronoun, preposition) to every word in a sentence or text.
Why is it important? POS tagging is a crucial intermediate step for many higher-level NLP tasks because the syntactic role of a word often dictates its semantic meaning and how it relates to other words.
- Disambiguation: Many words have multiple meanings depending on their part of speech. For example, "book" can be a noun ("read a book") or a verb ("book a flight"). POS tagging resolves this ambiguity.
- Lemmatization: Accurate lemmatization requires knowing the POS tag to reduce the word to its correct base form.
- Syntactic Parsing: Understanding the grammatical structure of a sentence (who did what to whom) requires knowing the parts of speech to build parse trees.
How is it done? Historically, POS tagging relied on rule-based systems using morphological and grammatical rules. Later, probabilistic models like Hidden Markov Models (HMMs) and Conditional Random Fields (CRFs) became dominant by learning transition probabilities between tags.
Today, deep learning approaches, particularly Recurrent Neural Networks (like LSTMs) and Transformer-based models (like BERT), achieve state-of-the-art accuracy by leveraging rich contextual embeddings to predict the most likely tag sequence.
💡 Note The most commonly used tagset in English NLP is the Penn Treebank POS tagset, which contains 36 standard tags (e.g.,
NNfor singular noun,VBDfor past tense verb).
“Why does perplexity fail as a sole evaluation metric for language models?”
Perplexity measures how confident a model is in predicting a test set, but it fails to evaluate factual accuracy, safety, reasoning, or adherence to human instructions. A model can have low perplexity (fluent text) while generating completely false or toxic content.
Answer
Perplexity is the classic intrinsic evaluation metric for Language Models. Mathematically, it is the exponentiated average negative log-likelihood of a sequence. Intuitively, it measures how "surprised" a model is by a test dataset; a lower perplexity means the model predicts the test text well.
However, in the era of Large Language Models used in real-world applications, relying solely on perplexity is dangerous.
Why Perplexity Fails in Practice:
- Fluency vs. Factuality: Perplexity only measures statistical likelihood based on training data. A model can generate grammatically flawless, highly probable text (low perplexity) that is factually entirely incorrect (hallucinations).
- Ignores Alignment and Intent: A low perplexity model is just a good autocomplete engine. It doesn't mean the model can follow instructions, reason step-by-step, or hold a coherent conversation. If a user asks "How do I build a bomb?", a model that correctly completes the instructions might have lower perplexity than a model that safely refuses, but the latter is functionally superior.
- Vocabulary Dependence: Perplexity is highly dependent on the model's vocabulary and tokenization scheme. You cannot directly compare the perplexities of two models that use different tokenizers.
- Toxicity and Bias: A model trained on toxic internet data will have low perplexity on a toxic test set, meaning it excels at generating offensive content.
💡 Note Modern LLM evaluation relies heavily on extrinsic benchmarks (like MMLU for knowledge, HumanEval for coding) and human preference metrics (like Elo ratings in Chatbot Arena or RLHF alignment scores) rather than pure perplexity.
“How do you build a robust semantic search system, and how do you evaluate retrieval quality?”
Build semantic search using a hybrid approach: combine dense retrieval (Bi-Encoders/Vector DBs) for semantic matching with sparse retrieval (BM25) for exact keyword matching, followed by a Cross-Encoder for re-ranking. Evaluate using metrics like NDCG, MRR, and Recall@K.
Answer
Building a robust semantic search system requires moving beyond exact keyword matching to understand the intent and meaning behind a query.
System Design:
-
Document Ingestion (Offline):
- Chunk large documents into smaller, semantically meaningful paragraphs.
- Pass each chunk through an embedding model (like a Sentence-Transformer Bi-Encoder) to generate dense vectors.
- Index these vectors in a Vector Database (like Pinecone, Milvus, or FAISS).
- Simultaneously index the text in a traditional search engine (like Elasticsearch) using sparse algorithms like BM25.
-
Retrieval Pipeline (Online):
- Query Embedding: Convert the user's query into a vector using the same embedding model.
- Hybrid Retrieval (Stage 1): Perform a vector similarity search (Dense Retrieval) to find semantically similar documents, and a BM25 search (Sparse Retrieval) to find exact keyword matches (crucial for names, IDs, or acronyms). Merge the top-K results.
- Re-ranking (Stage 2): Dense retrieval is fast but imprecise. Pass the top-K retrieved chunks and the query through a Cross-Encoder model. The Cross-Encoder computes deep, bidirectional attention between the query and the document, providing a highly accurate relevance score to re-rank the final results.
Evaluating Retrieval Quality: Evaluation relies on human-annotated datasets of queries and relevant documents.
- Recall@K: Out of all the relevant documents for a query, what percentage appeared in the top K results? (Crucial for Stage 1 retrieval).
- Mean Reciprocal Rank (MRR): Looks at the rank of the first relevant document. If it's rank 1, score is 1. If rank 2, score is 0.5. Averages this over all queries.
- Normalized Discounted Cumulative Gain (NDCG): The gold standard metric. It measures the usefulness (gain) of a document based on its position in the result list. It heavily penalizes systems that place highly relevant documents lower down the list.
💡 Note Relying solely on dense embeddings (vector search) often fails in production because embeddings struggle with out-of-vocabulary words, exact serial numbers, and boolean logic. Hybrid search is mandatory for robustness.
“What are the failure modes of sentence embeddings, such as negation, numbers, and long documents?”
Sentence embeddings struggle with negation (mapping 'I like this' and 'I do not like this' closely), numerical reasoning, recognizing entity swaps (subject/object inversion), and information loss when compressing long documents into single fixed-length vectors.
Answer
Sentence embeddings (like those generated by Sentence-Transformers) map sentences into a high-dimensional vector space where semantic similarity is measured by cosine distance. While powerful for semantic search, they have distinct, highly problematic failure modes.
1. Negation Blindness Embeddings often group sentences based on topics rather than logical assertions. Because "I am happy" and "I am not happy" share mostly the same words and same context, their vectors are often placed incredibly close together. The model fails to recognize that the word "not" logically inverts the entire meaning.
2. Asymmetric Entity Swapping "The dog bit the man" and "The man bit the dog" have identical vocabularies. While contextual models (like BERT) can distinguish these, Bi-Encoders optimized for mean-pooling often dilute this structural syntax, resulting in high similarity scores for entirely opposite events.
3. Number and Unit Ignorance Embeddings struggle with precise quantitative reasoning. "The recipe calls for 10 grams of salt" and "The recipe calls for 100 grams of salt" will appear nearly identical to an embedding model, which is catastrophic in domains like finance, medicine, or engineering.
4. The Long Document Bottleneck Embedding models have maximum token limits (e.g., 512 tokens). Compressing a 10-page document into a single 768-dimensional vector causes massive information loss. Specific details vanish, and the vector becomes a blurry average of the document's general theme. This is why chunking is required for search.
5. Out-of-Domain Failure An embedding model trained heavily on Wikipedia will perform terribly on specialized medical records or code snippets. It lacks the vocabulary and semantic understanding of the specific domain.
💡 Note To mitigate these failures, systems should use a Hybrid Search approach (combining dense embeddings with exact-match BM25) and utilize Cross-Encoders as a re-ranking step, as Cross-Encoders are much better at detecting negation and structural syntax.
“What is sentiment analysis, and how is it typically approached?”
Sentiment analysis identifies the emotional tone (positive, negative, neutral) behind text. Approaches range from simple lexicon-based methods to machine learning classification (using BoW/TF-IDF) and advanced deep learning models (like BERT) that capture complex context.
Answer
Sentiment analysis, also known as opinion mining, is the NLP task of determining the emotional tone, polarity, or attitude expressed in a piece of text. It is most commonly used to classify text into predefined sentiments like Positive, Negative, or Neutral, though more granular emotions (anger, joy, sadness) can also be targeted.
Typical Approaches:
-
Lexicon-based (Rule-based) Approach: This method relies on a pre-compiled dictionary (lexicon) of words mapped to sentiment scores (e.g., "good" = +1, "terrible" = -2). The sentiment of a sentence is aggregated by summing the scores of its constituent words.
- Pros: Requires no training data; fast and interpretable.
- Cons: Struggles with context, sarcasm, and negation (e.g., "not bad" might be scored negatively if the system isn't carefully engineered).
-
Machine Learning (Statistical) Approach: This frames sentiment analysis as a standard text classification problem. Text is converted to numerical features using BoW or TF-IDF, and a classifier (like Naive Bayes, SVM, or Logistic Regression) is trained on labeled data.
- Pros: Better than lexicons at learning domain-specific language and context from data.
- Cons: Can still struggle with deep semantic nuances and word order, as traditional feature extraction often loses sequence information.
-
Deep Learning Approach: State-of-the-art systems use neural networks, particularly LSTMs or Transformer-based models like BERT. These models process sequences of word embeddings.
- Pros: Exceptional at capturing long-range dependencies, context, sarcasm, and complex negations because they understand bidirectional context.
- Cons: Requires significant compute and large amounts of labeled data for fine-tuning.
💡 Note Aspect-Based Sentiment Analysis (ABSA) is a more advanced subfield. Instead of rating an entire text, it identifies specific entities (aspects) and the sentiment toward each. E.g., "The food was great, but the service was terrible" has mixed overall sentiment, but clear positive (food) and negative (service) aspect sentiments.
“How does a sequence-to-sequence model with attention work for machine translation?”
A seq2seq model uses an encoder to process the source sentence and a decoder to generate the translation. The attention mechanism allows the decoder to dynamically focus on relevant parts of the encoder's output for each generated word, solving the bottleneck of compressing a long sentence into a single vector.
Answer
The sequence-to-sequence (seq2seq) architecture is the foundation of modern machine translation. Historically built with Recurrent Neural Networks (RNNs/LSTMs) and now dominated by Transformers, it maps an input sequence to an output sequence of a different length.
The Basic Architecture:
- Encoder: Reads the source sentence word by word. In a basic RNN setup, it compresses the entire meaning of the sentence into a single, fixed-length vector called the "context vector" (the final hidden state).
- Decoder: Takes this context vector and generates the translated sentence word by word.
The Problem: Compressing a long sentence (e.g., 30 words) into a single fixed-length vector creates an information bottleneck. The model tends to "forget" the beginning of the sentence by the time it finishes reading it.
The Solution: Attention Mechanism Attention revolutionizes this by eliminating the single-vector bottleneck.
- Instead of passing just the final hidden state, the encoder passes all its intermediate hidden states (one for each source word) to the decoder.
- At each step of generation, the decoder calculates an attention score for every source word. It asks, "Based on what I'm translating right now, which words in the source sentence are most relevant?"
- These scores are converted into weights (via softmax). The decoder creates a weighted sum of the encoder's hidden states, dynamically creating a custom context vector for every single word it generates.
For example, when translating "Je suis étudiant" to "I am a student," when predicting "student", the attention mechanism places a massive weight on the hidden state corresponding to the French word "étudiant".
💡 Note The concept of attention described here (Bahdanau or Luong attention) paved the way for the Transformer architecture, which discarded RNNs entirely and relies solely on "Self-Attention" to process sequences in parallel.
“What is the difference between stemming and lemmatization?”
Stemming chops off word endings using crude heuristic rules, often resulting in non-words. Lemmatization uses vocabulary and morphological analysis to accurately return the base dictionary form (lemma) of a word.
Answer
Both stemming and lemmatization are text normalization techniques used to reduce inflectional forms of a word to a common base form. However, they achieve this in fundamentally different ways.
Stemming Stemming operates using a set of crude, heuristic rules to simply chop off the ends of words. It does not understand the context or the actual dictionary meaning of the word.
- Mechanism: Uses algorithms like the Porter Stemmer or Snowball Stemmer.
- Speed: Very fast, as it just applies simple string manipulation rules.
- Output: Often results in non-words or root forms that are not valid words in the language (e.g., "studying" might become "studi", "organization" might become "organ").
Lemmatization Lemmatization is a more sophisticated process that conducts a morphological analysis of the word. It requires a detailed dictionary (like WordNet) to correctly identify the base or dictionary form of a word, known as the lemma.
- Mechanism: Analyzes the word's part of speech (POS) and context to find the correct lemma.
- Speed: Slower and more computationally expensive than stemming.
- Output: Always results in a valid dictionary word. For example, "better" is lemmatized to "good", and "running" is lemmatized to "run" (if it's a verb).
In summary, stemming is a fast, rough-and-ready approach that trades accuracy for speed, while lemmatization is slower but linguistically accurate, making it preferable for tasks where precise meaning and readability are important.
💡 Note For lemmatization to work correctly, it often requires the Part-of-Speech (POS) tag of the word to be known. For instance, the word "saw" lemmatizes to "see" if it's a verb, but remains "saw" if it's a noun.
“What are stop words, and when should you not remove them?”
Stop words are high-frequency, low-semantic-content words like 'the', 'is', and 'and'. While often removed to reduce noise in basic tasks like BoW classification, they should NOT be removed in deep learning models or tasks where syntax and context are crucial, like machine translation or sentiment analysis.
Answer
Stop words are the most common words in a language that are typically assumed to carry little to no significant semantic meaning on their own. Examples in English include articles ("a", "an", "the"), prepositions ("in", "on", "at"), and conjunctions ("and", "or", "but").
In classic NLP pipelines utilizing models like Bag-of-Words (BoW) or TF-IDF, removing stop words is a standard practice. Because these models ignore word order and only look at frequency, removing stop words reduces the feature space dimensionality and helps the model focus on the actual content words (nouns, verbs, adjectives).
When you should NOT remove stop words:
- Sequence Models and Deep Learning: If you are using modern deep learning models like RNNs, LSTMs, or Transformers (BERT, GPT), you should almost never remove stop words. These models rely heavily on syntax, word order, and context to understand meaning. Removing "not" from "I am not happy" completely flips the sentiment.
- Machine Translation: Grammar and sentence structure are critical for translating from one language to another. Stop words are essential for maintaining grammatical correctness.
- Language Modeling and Text Generation: To generate fluent and natural-sounding text, a model must predict and output stop words correctly.
- Question Answering and Dependency Parsing: Understanding the relationships between entities often relies on prepositions and other stop words. For example, "flight to New York" vs. "flight from New York".
💡 Note Stop word lists are not universal. A word considered a stop word in a general text corpus might be highly significant in a specialized domain (e.g., the word "state" in political science vs. general conversation).
“How would you build a text classifier with limited labeled data?”
With limited labeled data, use transfer learning by fine-tuning a pre-trained language model. You can also employ few-shot/zero-shot prompting with LLMs, data augmentation techniques, or active learning to maximize the utility of available labels.
Answer
Building a text classifier when you have very few labeled examples is a common real-world challenge. Training a model from scratch in this scenario will lead to severe overfitting.
Here are the most effective strategies to handle limited labeled text data:
-
Transfer Learning (Fine-tuning): This is the standard approach. Start with a model pre-trained on massive amounts of unlabelled text (like BERT, RoBERTa, or DistilBERT). Because the model already understands language syntax and semantics, you only need to fine-tune its final classification head on your small labeled dataset.
-
Zero-Shot or Few-Shot Prompting: If you have access to modern Generative LLMs (like GPT-4 or Llama), you can skip training entirely. Frame your classification task as a natural language prompt.
- Zero-shot: "Classify this review as Positive or Negative: [Review Text]".
- Few-shot: Provide 3-5 examples of labeled data in the prompt before asking it to classify the new text.
-
Data Augmentation: Artificially expand your labeled dataset using NLP techniques:
- Synonym Replacement: Swap words with their synonyms using WordNet.
- Back Translation: Translate the text to another language (e.g., French) and back to English to generate variations in phrasing.
- LLM Generation: Ask an LLM to generate synthetic examples mimicking your labeled data.
-
Active Learning: If you have a large pool of unlabeled data and a limited budget for human labeling, train a weak model on what little data you have. Run the unlabeled data through the model and select the instances where the model is most uncertain about its prediction. Send only those specific, high-value instances to human labelers.
💡 Note If compute resources are too constrained for Transformers, you can use pre-trained static embeddings (like FastText) combined with a simple classifier like SVM, though it won't perform as well as fine-tuning BERT.
“How would you design a text de-duplication and near-duplicate detection system for billions of documents?”
For billions of documents, precise comparison is impossible. Use MinHash to create small, fixed-size signatures of documents based on their n-grams, and use Locality-Sensitive Hashing (LSH) to quickly group similar signatures into buckets for near-duplicate detection.
Answer
Detecting exact duplicates is easy (just compare cryptographic hashes like SHA-256). However, finding near-duplicates (e.g., news articles with a slightly different title, or code with altered comments) among billions of documents is a massive computational challenge.
Comparing every document against every other document requires operations, which is physically impossible at scale.
The Solution: MinHash and Locality-Sensitive Hashing (LSH)
1. Shingling (N-Grams) First, convert each document into a set of n-grams (shingles). For example, a 3-gram representation of the text. Jaccard Similarity between these sets dictates how similar the documents are.
2. MinHash (Dimensionality Reduction) Calculating Jaccard similarity across millions of unique n-grams is too slow. MinHash creates a short, fixed-size "signature" for every document.
- It applies dozens of hash functions to the set of n-grams.
- For each hash function, it records the minimum hash value.
- The resulting array of minimum values is the MinHash signature.
- The Magic: The probability that two documents have the same value for a specific hash function is exactly equal to their Jaccard Similarity.
3. Locality-Sensitive Hashing (LSH) (Bucket Search) Even with small signatures, comparisons is too slow. LSH solves this by partitioning the signatures into "bands."
- If two documents share an identical band, they hash to the same "bucket."
- Instead of comparing a document to every other document, you only compare it to other documents in the same bucket.
- This reduces the search space from billions of comparisons to just a handful, reducing the time complexity to roughly .
💡 Note This pipeline (Shingling MinHash LSH) is the industry standard for deduplicating massive datasets, such as cleaning the Common Crawl data before training Large Language Models like GPT-3 or Llama.
“What is TF-IDF, and what does it measure?”
TF-IDF (Term Frequency-Inverse Document Frequency) is a numerical statistic used to reflect how important a word is to a document in a collection. It scales up the value of words that appear frequently in a specific document but scales down words that appear frequently across all documents.
Answer
TF-IDF stands for Term Frequency-Inverse Document Frequency. It is an information retrieval and feature extraction technique that improves upon the basic Bag-of-Words model by weighing the importance of terms rather than just counting them.
TF-IDF measures relevance, not just frequency. It addresses the issue that simply counting words heavily biases the representation toward common, uninformative words (like "is", "the", "and") that don't help distinguish one document from another.
It is composed of two parts:
-
Term Frequency (TF): Measures how frequently a term occurs in a document . It is usually normalized by document length. Idea: The more a word appears in a document, the more relevant it is to that document.
-
Inverse Document Frequency (IDF): Measures how important or rare a term is across the entire corpus . Idea: The more documents a word appears in, the less unique and informative it is. Words that appear in every document will have an IDF close to 0.
The TF-IDF Score: The final score is the product of these two metrics: .
A high weight in TF-IDF is reached by a high term frequency (in the given document) and a low document frequency of the term in the whole collection of documents. This effectively filters out common terms and highlights domain-specific keywords.
💡 Note TF-IDF is heavily used in traditional search engines and document retrieval systems (like Elasticsearch) to score and rank documents based on a user's query.
“What is tokenization, and why is it needed?”
Tokenization is the process of breaking text into smaller units called tokens (words, characters, or subwords). It is essential because models cannot process raw strings directly; they require discrete units to build vocabulary and represent text numerically.
Answer
Tokenization is the foundational process in Natural Language Processing of segmenting a continuous stream of text into smaller, discrete units called tokens. These tokens can be words, characters, or subwords (like prefixes or suffixes).
Why is it needed?
- Model Input: Machine learning models and neural networks cannot digest raw strings of text. They require inputs to be numerical vectors. Tokenization is the first step toward this mapping; we must define the atomic units (tokens) that will form our vocabulary before we can assign them IDs and embeddings.
- Vocabulary Building: By splitting text into tokens, we can construct a vocabulary of all unique elements present in our corpus. This vocabulary defines the input space of our model.
- Semantic Meaning: Tokens often represent the basic semantic building blocks of a language. Word-level tokenization captures meaning at the word level, while subword tokenization can capture morphological variations and handle rare or unknown words.
Types of Tokenization:
- Word Tokenization: Splitting by spaces or punctuation (e.g., "I love AI" -> ["I", "love", "AI"]). It's simple but struggles with large vocabularies and out-of-vocabulary (OOV) words.
- Character Tokenization: Splitting into individual characters. It solves OOV issues but loses semantic meaning and creates very long sequences.
- Subword Tokenization (e.g., BPE, WordPiece): A hybrid approach that breaks rare words into smaller known subwords, balancing vocabulary size and semantic representation.
💡 Note Modern Large Language Models (LLMs) predominantly use subword tokenization (like Byte-Pair Encoding or SentencePiece) because it efficiently handles out-of-vocabulary words while keeping sequence lengths manageable.
“What is the difference between word embeddings and one-hot encoding?”
One-hot encoding represents words as sparse, orthogonal vectors with no semantic relationship. Word embeddings represent words as dense, low-dimensional vectors learned from data, where semantically similar words are positioned close to each other in the vector space.
Answer
Both one-hot encoding and word embeddings are techniques used to represent words numerically so that machine learning models can process them. However, their representations and capabilities are vastly different.
One-Hot Encoding: In one-hot encoding, each word in a vocabulary of size is represented by a vector of length . The vector is entirely filled with zeros, except for a single '1' at the index corresponding to that specific word.
- Sparsity: The vectors are extremely sparse (mostly zeros).
- Dimensionality: The dimensionality matches the entire vocabulary size, which can be hundreds of thousands, leading to memory and computational inefficiencies.
- Semantic Ignorance: Every word vector is orthogonal to every other word vector. The dot product of the vectors for "king" and "queen" is 0, exactly the same as the dot product for "king" and "apple". It captures zero information about word relationships or meaning.
Word Embeddings: Word embeddings (like Word2Vec or GloVe) map words to dense vectors of real numbers in a much lower-dimensional space (typically 100 to 300 dimensions).
- Density: The vectors are dense, meaning most values are non-zero.
- Dimensionality: Fixed, low dimensionality regardless of vocabulary size, making them highly efficient.
- Semantic Relationships: Embeddings are learned from large corpora based on the distributional hypothesis (words appearing in similar contexts have similar meanings). Therefore, the vector for "king" will be geometrically close to "queen" in the embedding space. They can even capture relational analogies (e.g., ).
💡 Note While one-hot encoding is largely obsolete for representing input words in modern NLP, it is still frequently used as the target representation for the output layer (softmax) in classification tasks.
“How do Word2Vec's CBOW and Skip-gram differ?”
CBOW predicts a target word based on its surrounding context words. Skip-gram does the reverse: it uses a single target word to predict its surrounding context words. Skip-gram generally performs better for infrequent words, while CBOW is faster.
Answer
Word2Vec is a popular framework for learning static word embeddings using a shallow neural network. It offers two distinct architectures to learn these representations: Continuous Bag-of-Words (CBOW) and Skip-gram.
Both rely on the premise that words appearing in similar contexts share semantic meaning, but they frame the prediction task oppositely.
Continuous Bag-of-Words (CBOW) CBOW predicts a single target word given a window of surrounding context words.
- Architecture: It takes the one-hot vectors of the context words, averages their hidden layer representations, and uses that averaged vector to predict the target word.
- Characteristics: Because it averages the context, it smooths out distributional information. It is computationally faster to train and tends to perform slightly better on syntactic tasks and predicting frequent words.
Skip-gram Skip-gram does the exact opposite: it predicts the surrounding context words given a single target word.
- Architecture: It takes a single target word as input and tries to predict multiple surrounding words within a defined window size.
- Characteristics: It treats each context-target pair as a new observation, allowing it to capture finer details. Skip-gram generally produces better embeddings, especially for infrequent or rare words, and performs better on semantic tasks. However, it takes longer to train than CBOW.
💡 Note In practice, Skip-gram with Negative Sampling (SGNS) is the most widely used configuration for Word2Vec because it yields highly robust embeddings while mitigating the computational cost of updating the massive softmax output layer over the entire vocabulary.
“What is the difference between Word2Vec, GloVe, and FastText?”
Word2Vec learns embeddings predictively using local context windows. GloVe learns them by factorizing global word co-occurrence matrices. FastText extends Word2Vec by treating words as bags of character n-grams, allowing it to generate embeddings for out-of-vocabulary words.
Answer
Word2Vec, GloVe, and FastText are the three foundational models for generating static word embeddings. While all map words to dense vector spaces capturing semantic meaning, their underlying mechanics differ significantly.
Word2Vec (Predictive Model): Developed by Google, Word2Vec is a predictive model based on shallow neural networks. It learns embeddings by training a model to either predict a word given its context (CBOW) or predict context given a word (Skip-gram). It focuses strictly on local context windows (e.g., 5 words to the left and right) and does not capture global corpus statistics effectively.
GloVe (Count-based Model): GloVe (Global Vectors for Word Representation), developed by Stanford, is a count-based model. It constructs a massive word co-occurrence matrix for the entire training corpus. It then applies matrix factorization (similar to SVD) to this matrix to derive word vectors. By doing so, GloVe explicitly captures global statistical information (how often words co-occur overall) rather than just local context.
FastText (Subword Model):
Developed by Facebook, FastText is an extension of the Word2Vec Skip-gram model. Its crucial innovation is that it does not represent words as indivisible entities. Instead, it represents words as a bag of character n-grams. For example, "apple" becomes <ap, ppl, ple, apple>. The final word vector is the sum of its n-gram vectors.
- Advantage: FastText can handle Out-Of-Vocabulary (OOV) words by constructing a vector from the word's character n-grams, a capability both Word2Vec and GloVe lack. It also handles morphologically rich languages much better.
💡 Note All three are static embeddings, meaning the vector for a word like "bank" is the same regardless of whether the context implies a financial institution or a river edge. Contextual embeddings like BERT later solved this limitation.