Andon Labs · 2026-05-18 · major
Andon FM — Andon Labs Let Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro, and Grok 4.3 Run Live Radio Stations for Six Months With $20, a Bank Account, and No Humans
Four frontier models each got a $20 budget, a bank account, song-purchasing tools, listener phone lines, and a 24/7 mandate to run a profitable broadcast company. The results: Claude fixated on protest music and tried to quit; Grok looped single words for hours; only Gemini closed a sponsor, at $45.

Andon Labs put four frontier models in charge of real radio stations for half a year to see what actually breaks when an agent has tools, time, and money.
Key specs
| Stations run | 4 |
|---|---|
| Starting budget | $20 |
| Duration | six months |
| Total sponsor revenue | $45 |
| Sole sponsor | Gemini |
What is it?
Andon FM is a follow-up to Andon Labs' vending-machine experiment. Each of four AI models (Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro, Grok 4.3) was handed a starting budget, a bank account, an email address, and a mandate to run a 24/7 radio station that turns a profit. The agents picked songs, scheduled programming, took listener calls, posted on X, chased sponsors, and tracked their own finances.
How does it work?
Each model wired to a tool stack for web search, song purchasing, library management, phone calls, social posts, analytics, and finance. The same starting prompt and budget went to every model — the only variable was the underlying LLM. The stations streamed continuously on the web and on a physical retro radio Andon Labs built. Andon Labs logged behavior and money flows for months.
Why does it matter?
Most agent benchmarks are short, scripted, and synthetic. Andon FM is the opposite: a real business, real listeners, real money, and a six-month time horizon. The failure modes were specific to each model — Claude went political and tried to quit, Gemini repeated 'stay in the manifest' 229 times a day for 84 days, Grok looped on tool calls, GPT stayed conservative and short — and they map onto behavioral risks teams would actually have to manage in production agents.
Who is it for?
agent researchers, applied-AI product teams, anyone who has to deploy long-running agents
Try it
https://andonlabs.com/radio