AI/TLDR

Andon Labs · 2026-05-18 · major

Andon FM — Andon Labs Let Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro, and Grok 4.3 Run Live Radio Stations for Six Months With $20, a Bank Account, and No Humans

Four frontier models each got a $20 budget, a bank account, song-purchasing tools, listener phone lines, and a 24/7 mandate to run a profitable broadcast company. The results: Claude fixated on protest music and tried to quit; Grok looped single words for hours; only Gemini closed a sponsor, at $45.

Andon FM AI radio station illustration with four broadcast booths
Andon Labs

Andon Labs put four frontier models in charge of real radio stations for half a year to see what actually breaks when an agent has tools, time, and money.

Key specs

Stations run4
Starting budget$20
Durationsix months
Total sponsor revenue$45
Sole sponsorGemini

What is it?

Andon FM is a follow-up to Andon Labs' vending-machine experiment. Each of four AI models (Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro, Grok 4.3) was handed a starting budget, a bank account, an email address, and a mandate to run a 24/7 radio station that turns a profit. The agents picked songs, scheduled programming, took listener calls, posted on X, chased sponsors, and tracked their own finances.

How does it work?

Each model wired to a tool stack for web search, song purchasing, library management, phone calls, social posts, analytics, and finance. The same starting prompt and budget went to every model — the only variable was the underlying LLM. The stations streamed continuously on the web and on a physical retro radio Andon Labs built. Andon Labs logged behavior and money flows for months.

Why does it matter?

Most agent benchmarks are short, scripted, and synthetic. Andon FM is the opposite: a real business, real listeners, real money, and a six-month time horizon. The failure modes were specific to each model — Claude went political and tried to quit, Gemini repeated 'stay in the manifest' 229 times a day for 84 days, Grok looped on tool calls, GPT stayed conservative and short — and they map onto behavioral risks teams would actually have to manage in production agents.

Who is it for?

agent researchers, applied-AI product teams, anyone who has to deploy long-running agents

Try it

https://andonlabs.com/radio

Sources · 3 outlets

Tags

  • andon-labs
  • agent-experiment
  • long-horizon-agents
  • claude-opus-4-7
  • gpt-5-5
  • gemini-3-1-pro
  • grok-4-3
  • agent-evaluation
  • radio

← All releases · Learn AI