AI/TLDR

SYSTEM 11/14 · THE FIELD GUIDE

Multimodal AI

Beyond text — models that see, hear, speak, draw, and film.

4 TRACKS55 ARTICLESbeginner → advanced

Vision & Document Understanding

How models read images, screenshots, documents, and video.

OPEN TRACK

Speech & Voice

Whisper-style transcription, neural voices, and the realtime voice agent stack.

OPEN TRACK

Image Generation

Diffusion, prompting for pixels, and the open image stack.

OPEN TRACK

Video, Audio & Beyond

The frontier modalities: video, world models, music, 3D, and any-to-any.

OPEN TRACK