Two Minute Papers · 2026-06-16 · notable
Two Minute Papers: 'They Looked Inside Claude's AI's Mind. It Got Weird'
Károly Zsolnai-Fehér's new Two Minute Papers video walks through the latest mechanistic interpretability work on Claude — what Anthropic's researchers found by probing the model's internal features.

A fresh Two Minute Papers explainer on Anthropic's mechanistic interpretability work inside Claude.
What is it?
A new Two Minute Papers video on YouTube. Károly Zsolnai-Fehér summarizes interpretability research that probes Claude's internal features — what neurons and circuits activate, and what surprised the researchers along the way.
How does it work?
The video follows the channel's signature short-form format: paper figures, animated explanations, and on-screen captions over a few minutes. Zsolnai-Fehér frames the findings as a non-researcher introduction to mechanistic interpretability work on a frontier model.
Why does it matter?
Two Minute Papers is one of the largest research-explainer channels on YouTube. A fresh explainer like this is often the first place a broad audience hears about interpretability research, and it shapes how the public reads Anthropic's safety positioning during the current Fable 5 / Mythos 5 suspension news cycle.
Frequently asked questions
- What is the Two Minute Papers video about?
- The Two Minute Papers video, titled 'They Looked Inside Claude's AI's Mind. It Got Weird', summarizes mechanistic interpretability research on Anthropic's Claude. Host Károly Zsolnai-Fehér walks through what researchers found by probing the model's internal features — which neurons and circuits activate inside Claude, and what surprised the team along the way.
- Who makes the Two Minute Papers channel?
- Two Minute Papers is hosted by Károly Zsolnai-Fehér on YouTube. It is one of the largest research-explainer channels on the platform, known for a short-form format that uses paper figures, animated explanations, and on-screen captions to introduce AI and graphics research to a broad, non-researcher audience.
- What is mechanistic interpretability in this context?
- In this video, mechanistic interpretability refers to research that probes Claude's internal features to understand how the model works. Zsolnai-Fehér frames it as a non-researcher introduction to the work, describing what neurons and circuits activate inside the frontier model and what the Anthropic researchers found surprising.
- Where can I watch the video?
- The Two Minute Papers video 'They Looked Inside Claude's AI's Mind. It Got Weird' is available on YouTube. It follows the channel's signature short format, running a few minutes and pairing animated explanations with on-screen captions to summarize the interpretability research on Claude.