Evaluation & Safety · TRACK 03/05
Benchmarks & Leaderboards
MMLU to SWE-bench to LMArena — what the scores mean and when they lie.
// THE TRACK
01 · START HERELLM BenchmarksUnderstand what benchmark scores like MMLU and GPQA actually measure and how to read them on a model announcement.BEGINNERChatbot Arena (LMSYS)Understand how Chatbot Arena turns millions of human votes into a live AI leaderboard — and why that is both more realistic and more gameable than a fixed test.BEGINNERCoding BenchmarksUnderstand how coding ability is benchmarked, from single-function tests to fixing real GitHub issues.BEGINNERLMArena & Elo RatingsUnderstand how LMArena turns blind human votes into Elo-style rankings — and the caveats behind the leaderboard.BEGINNERAgent BenchmarksUnderstand how agent benchmarks score end-to-end task completion instead of single answers.INTERMEDIATEBenchmark ContaminationUnderstand how test data leaks into training sets, how researchers detect it, and why it quietly inflates scores.INTERMEDIATEBenchmark OverfittingYou'll understand how chasing a benchmark can inflate scores while real-world ability stagnates, and how that differs from contamination.INTERMEDIATEPublic vs Private BenchmarksYou'll understand why a model topping public leaderboards may still fail your task, and why teams keep private, held-out benchmarks.BEGINNERReading a Benchmark ScoreYou'll understand what a benchmark number actually tells you and the hidden settings that can make the same model look very different.BEGINNERMMLUYou will understand what MMLU measures, why it shaped benchmark culture, and why top models have largely saturated it.INTERMEDIATEGPQAYou will understand what makes GPQA 'Google-proof' and why its Diamond subset became a key reasoning benchmark.INTERMEDIATESWE-benchYou will understand how SWE-bench evaluates coding agents by checking real patches against a repo's test suite.INTERMEDIATEARC-AGIYou will understand why ARC-AGI tests abstract reasoning rather than memorized knowledge and why it resists scaling.INTERMEDIATEArtificial AnalysisYou will understand how Artificial Analysis combines quality, speed, and price into one comparison view and how to interpret it.INTERMEDIATEStanford HELMYou will understand why HELM evaluates models across many aspects at once and how that differs from a single-score benchmark.INTERMEDIATE