█

AI/TLDR

Two Minute Papers · 2026-10-11 · notable

Two Minute Papers — 'DeepSeek Solved AI’s Memory Problem'

Two Minute Papers covers DeepSeek-OCR, DeepSeek's open MIT-licensed model that packs a page of text into as few as 64 vision tokens, an idea DeepSeek calls contexts optical compression.

Two Minute Papers video thumbnail for the DeepSeek-OCR episode

Károly Zsolnai-Fehér explains DeepSeek's idea of storing long text as images to save tokens.

What is it?

'DeepSeek Solved AI’s Memory Problem' went up on Two Minute Papers on 11 October 2026. The video description names the DeepSeek-OCR GitHub repository as its source paper.

How does it work?

DeepSeek-OCR renders text as an image and reads it back through a vision encoder, so a page costs a fixed number of vision tokens: 64 at 512×512, 100 at 640×640, 256 at 1024×1024 and 400 at 1280×1280. DeepSeek reports about 2,500 tokens per second for PDF processing on one A100-40G with vLLM.

Why does it matter?

Long contexts are expensive because every text token costs memory and compute. Compressing old context into a small number of vision tokens is one way to let a model keep more history for less, which is the 'memory problem' the video title refers to.

Who is it for?

anyone curious about long-context LLMs

Sources · 2 outlets

Tags

  • video
  • two-minute-papers
  • deepseek
  • deepseek-ocr
  • ocr
  • vision-language
  • context-compression

← All releases · Learn AI