AI/TLDR

SYSTEM 13/14 · THE FIELD GUIDE

Evaluation & Safety

Measuring whether models are good — and keeping them from being bad.

5 TRACKS63 ARTICLESbeginner → advanced

Evaluation Basics

Because "it looks good to me" is not a test suite.

OPEN TRACK

LLM-as-a-Judge

Using models to grade models — and when to distrust the grader.

OPEN TRACK

Benchmarks & Leaderboards

MMLU to SWE-bench to LMArena — what the scores mean and when they lie.

OPEN TRACK

Red Teaming & Jailbreaks

Attack your own AI before someone else does.

OPEN TRACK

Alignment & Safety Basics

Why models refuse, how they're steered, and the bigger risk map.

OPEN TRACK