LLM-as-a-Judge
An automated evaluation paradigm where a powerful language model is prompted to score, critique, or rank the outputs of other AI models.
Think of It Like This
Like having a master chef taste and grade the dishes prepared by culinary students during a cooking exam.
This technique scales evaluation far beyond what human annotators can manage, enabling rapid iterations during model development. The judge LLM is provided with a strict grading rubric and often asked to generate a chain-of-thought rationale before assigning a final score. However, it must be carefully monitored for inherent biases.