AI/TLDR

Two Minute Papers · 2026-06-16 · notable

Two Minute Papers: 'They Looked Inside Claude's AI's Mind. It Got Weird'

Károly Zsolnai-Fehér's new Two Minute Papers video walks through the latest mechanistic interpretability work on Claude — what Anthropic's researchers found by probing the model's internal features.

Two Minute Papers thumbnail for the video on Claude interpretability research.

A fresh Two Minute Papers explainer on Anthropic's mechanistic interpretability work inside Claude.

What is it?

A new Two Minute Papers video on YouTube. Károly Zsolnai-Fehér summarizes interpretability research that probes Claude's internal features — what neurons and circuits activate, and what surprised the researchers along the way.

How does it work?

The video follows the channel's signature short-form format: paper figures, animated explanations, and on-screen captions over a few minutes. Zsolnai-Fehér frames the findings as a non-researcher introduction to mechanistic interpretability work on a frontier model.

Why does it matter?

Two Minute Papers is one of the largest research-explainer channels on YouTube. A fresh explainer like this is often the first place a broad audience hears about interpretability research, and it shapes how the public reads Anthropic's safety positioning during the current Fable 5 / Mythos 5 suspension news cycle.

Frequently asked questions

What is the Two Minute Papers video about?
The Two Minute Papers video, titled 'They Looked Inside Claude's AI's Mind. It Got Weird', summarizes mechanistic interpretability research on Anthropic's Claude. Host Károly Zsolnai-Fehér walks through what researchers found by probing the model's internal features — which neurons and circuits activate inside Claude, and what surprised the team along the way.
Who makes the Two Minute Papers channel?
Two Minute Papers is hosted by Károly Zsolnai-Fehér on YouTube. It is one of the largest research-explainer channels on the platform, known for a short-form format that uses paper figures, animated explanations, and on-screen captions to introduce AI and graphics research to a broad, non-researcher audience.
What is mechanistic interpretability in this context?
In this video, mechanistic interpretability refers to research that probes Claude's internal features to understand how the model works. Zsolnai-Fehér frames it as a non-researcher introduction to the work, describing what neurons and circuits activate inside the frontier model and what the Anthropic researchers found surprising.
Where can I watch the video?
The Two Minute Papers video 'They Looked Inside Claude's AI's Mind. It Got Weird' is available on YouTube. It follows the channel's signature short format, running a few minutes and pairing animated explanations with on-screen captions to summarize the interpretability research on Claude.

Sources

Tags

  • video
  • two-minute-papers
  • anthropic
  • claude
  • interpretability
  • explainer

← All releases · Learn AI