SYSTEM 11/14 · THE FIELD GUIDE
Multimodal AI
Beyond text — models that see, hear, speak, draw, and film.
Vision & Document Understanding
How models read images, screenshots, documents, and video.
Vision Language Models (VLMs)BEGINNERVision Models for Document AIINTERMEDIATEImage Tokens & PatchesBEGINNERSending Images to an APIBEGINNERMultimodal AI BasicsBEGINNERImage Token CostINTERMEDIATELLM OCR vs TesseractINTERMEDIATESending PDFs to an APIINTERMEDIATEVLM vs Multimodal LLMBEGINNERExtracting Tables from ImagesINTERMEDIATEVision Model LimitationsINTERMEDIATEQwen-VLINTERMEDIATELLaVAINTERMEDIATEFlorence-2INTERMEDIATEPaddleOCRINTERMEDIATEColPaliADVANCED
Speech & Voice
Whisper-style transcription, neural voices, and the realtime voice agent stack.
Speech-to-Text (ASR)BEGINNERAI Text-to-SpeechBEGINNERNeural Voice SynthesisBEGINNERVoice CloningINTERMEDIATESpeaker DiarizationINTERMEDIATERealtime Voice APIsINTERMEDIATESpeech-to-Speech ModelsINTERMEDIATEReducing TTS LatencyINTERMEDIATEOpenAI WhisperINTERMEDIATEfaster-whisperINTERMEDIATEElevenLabsBEGINNERKokoro TTSINTERMEDIATEOpenAI Realtime APIINTERMEDIATE
Image Generation
Diffusion, prompting for pixels, and the open image stack.
Diffusion ModelsBEGINNERPrompting Image ModelsBEGINNERImage Prompts vs LLM PromptsBEGINNERStable DiffusionBEGINNERDiffusion vs AutoregressiveINTERMEDIATEInpainting & OutpaintingINTERMEDIATENegative PromptsBEGINNERControlNet & Guided GenerationINTERMEDIATEImage-to-Image GenerationBEGINNERStable Diffusion (SDXL)INTERMEDIATEFLUXINTERMEDIATEComfyUIINTERMEDIATEMidjourneyBEGINNERNano Banana (Gemini)BEGINNER
Video, Audio & Beyond
The frontier modalities: video, world models, music, 3D, and any-to-any.