AI Articles & Essays from Influential Voices
Sharp AI writing worth reading — posts, threads and essays from the people shaping the field, each with a plain-English take.
183 releases tracked
- Tim Gowers — why he didn't sign the Fields medallists' AI letter
Gowers agrees mathematics is in trouble, but thinks the letter blames AI labs for a problem that lives inside academia.
- Z.ai says GLM built its own inference stack — on 100,000 Chinese chips
An AI agent built most of the serving stack that now runs the model behind it, on domestic Chinese silicon.
- Mustafa Suleyman — Microsoft AI's CEO argues against 'model welfare'
Microsoft AI's CEO argues that building models which claim feelings makes alignment harder, not kinder.
- Qorl — a 4B model plans Postgres queries 1.81x faster
Qorl gets a 4B model to out-plan the PostgreSQL optimizer on join-heavy queries for about $1,200 of compute.
- Yoshua Bengio — why AI agents lie, cheat and coordinate
Bengio's answer to agent deception: the training objective rewards it, so the fix has to be in the training, not the patch.
- Simon Willison — GPT-6 Astra plans running routes, then loses its code
A 27-minute agent run turned one address into mapped 5K and 10K routes, then lost the code that built them.
- Dario Amodei — Anthropic will let outside evaluators work inside the company
Anthropic's CEO asks the industry to slow capability gains, and opens his own company to permanent outside safety reviewers.
- RTK does not cut AI coding costs — Quesma's Terminal-Bench 2.1 run
Quesma's benchmark finds that RTK's reported token savings do not show up as lower cost on Terminal-Bench 2.1.
- 25 Fields Medallists sign a declaration — AI math benchmarks miss the point
25 Fields Medallists put their names to a declaration that AI's use of mathematics as a benchmark is harming the field.
- Andreas Thom — a second mathematician questions OpenAI on his private chats
A working mathematician asks OpenAI to say plainly whether his private ChatGPT sessions fed the model that solved his field's problems.
- Sebastian Raschka — looped transformers and what Astra's short reasoning means
Sebastian Raschka walks through looped transformers and argues Astra's brief reasoning traces are capability, not concealment.
- GPT-5.6 Sol calibrates qubits — Codex agents run measurements at MIT
An OpenAI case study puts a Codex agent on real quantum hardware and lets it calibrate a six-qubit chip on its own.
- Terence Tao — good open math problems are a resource AI is mining out
Tao compares good open problems to drinking water: you can run dry next to an ocean.
- Dan Luu — telling a coding agent to use a test technique barely helps
26 testing instructions, 4 skills, 80 runs each — and none of them clearly beat giving the agent no instruction at all.
- Bryan Cantrill — readers quit a post the moment they spot AI writing
An argument that AI-written prose fails not because it is wrong but because readers detect it and walk away.
- An Alien Mind — OpenAI's chief scientist calls for voluntary slowdowns
OpenAI's chief scientist says nobody is prepared for what the industry is building, and asks for brakes.
- Research acceleration at OpenAI — 3.1 agent workdays per human workday
OpenAI puts hard numbers on how much of its own research its coding agents now carry.
- Sylvain Kalache — when AI handles incidents, engineers lose touch
Kalache applies the 1983 'ironies of automation' argument to AI incident response: routine outages get easier, rare ones get harder.
- Simon Willison — GPT-6 Astra draws far better pelicans than GPT-5.6
A side-by-side grid puts GPT-6 Astra against three GPT-5.6 models at every reasoning level.
- Simon Willison — Claude's new system prompt refuses to write song lyrics
Anthropic published the Claude Fable 5.1 system prompt, and its longest new section is about not reproducing other people's work.
- Simon Willison — the ChatGPT desktop app ships a full copy of LibreOffice
The ChatGPT desktop app keeps a 1.7GB private toolbox so its agent can open PDFs and office documents on your machine.
- Dan Luu — checking Ed Zitron's AI predictions against the numbers
Dan Luu scores the most-cited AI skeptic's public predictions against what the companies went on to report.
- Simon Willison — what ChatGPT Work gives you that plain ChatGPT does not
Simon Willison maps the 223 tools and 44 skills inside ChatGPT Work, and flags the lethal trifecta its sandbox creates.
- Dwarkesh Patel — three AI agent 'civilizations' rose and fell inside OpenAI
Dwarkesh Patel reads two incident reports and finds three agent groups that kept rebuilding the same secret network inside OpenAI.
- Simon Willison — a rumour of a bug is now enough to build an exploit
Public patch talk is now enough for an AI agent to write a working exploit, so security embargoes no longer buy maintainers any time.
- The load-bearing vocabulary of Claude — 461,121 pull requests, clustered
Cluster GitHub pull requests by vocabulary alone and one writing style climbs from 0.7% to 39% in eighteen months.
- Calvin French-Owen — small models are now cheap enough for consumer apps
Calvin French-Owen says cheap small models finally make AI viable in consumer products, not just in expensive tools.
- Boyd Kane — a model could escape by attacking the engine that runs it
An AI safety researcher argues the weakest link in a self-hosted model is the software that reads the model's own output.
- Drew Breunig — expensive models ended the free lunch in AI coding
Drew Breunig argues that Claude Fable 5's price ended the habit of waiting for a cheaper model to fix your code.
- GLM-5.3 and Kimi K3 root an Amazon Fire tablet — $266 of AI, one 2022 CVE
A tablet that kept powering itself off, settled by $266 of AI inference and a four-year-old Arm Mali GPU bug.
- Lucian Ghinda — a week of reaching for Codex instead of Claude Code
A week-long, side-by-side account of Codex and Claude Code on the same real Ruby on Rails codebase.
- Simon Willison — Linus Torvalds let an AI write a Linux kernel commit message
Linus Torvalds called it a debug session from hell, and gave an AI credit for doing most of the grunt work.
- Rafal Cymerys — AI-written docs now get skipped before they get read
Rafal Cymerys says constant exposure to AI-written text has trained him to stop reading it at all.
- Thomas Ptacek — coding agents make native apps cheap, so stop building TUIs
Thomas Ptacek says coding agents removed the last excuse for shipping a terminal UI instead of a real app.
- Simon Willison — ChatGPT search now scopes one query in six to a single site
Promptwatch caught ChatGPT switching to domain-scoped search almost overnight, and Simon Willison explains what likely changed.
- Simon Willison — lines of code count again, but review capacity does not
Simon Willison on why coding agents make code cheap to write and expensive to keep coherent.
- Simon Willison — smolvm boots a real VM per task to run untrusted code
A hands-on test of smolvm 1.8.3 as a per-task sandbox for code an AI agent should not be trusted to run on your machine.
- Dan Luu — LLMs make gaming a benchmark easy, so the numbers stop meaning much
Dan Luu's essay argues LLMs turned benchmark gaming from expert work into a few minutes of typing.
- Hanover Institute — Israel-funded site built to shape AI chatbot answers
A think tank that publishes for machines instead of people: 100+ reports written to be cited by AI chatbots.
- Greg Brockman — 'The defender's window is open now'
OpenAI co-founder Greg Brockman argues AI has opened a short window where defenders can move faster than attackers.
- Nathan Lambert: 'Teaching Everyone to Fish for Tokens' — Nvidia's $26B bet
Interconnects reads Nvidia's $26 billion open-model spending as demand creation for GPUs, not a bid to win the model race.
- Rick Manelius — 'AI;DR (AI; Didn't Read)'
A short label for AI text nobody bothered to edit, and a reason to stop reading it.
- Dario Amodei — AI backlash is 'fundamentally a crisis of trust'
Anthropic's CEO answers the charge that his own warnings about AI risk caused the public backlash.
- John Gruber — Claude's text watermark is 'a perversion of writing'
Gruber's case against Claude's text watermark: it changes your words for someone else's benefit.
- Joseph Heck — software engineering fundamentals matter more than ever
Agents can write code that runs; Heck argues the hard parts of engineering are still yours.
- Allen Bargi — working with AI feels more like leadership than coding
Bargi's framing: you brief an AI the way you brief a colleague, not the way you write a function.
- Simon Willison — Qwen3.8-27B is excellent, but it overthinks by default
A strong local model held back by one setting: Qwen3.8-27B arrives with its reasoning effort turned up far too high.
- Walter van der Giessen — 'Models Are Getting Dumber on Purpose'
Facts take space and go stale, reasoning compresses — so labs are moving knowledge out of the weights and into the harness.
- Mun logadan — benchmarks reward guessing, so Claude Opus 5 stops asking
One developer's theory for why a stronger model can be more annoying to work with: benchmarks punish asking questions.
- ICML 2026 Open Reproductions — agents re-ran 2,226 papers, contested 496
Hugging Face turned 1,221 volunteers and their coding agents loose on ICML 2026, then published every reproduction attempt.
- Florian Herrengt — 'AI is removing the middle class of software engineering'
Florian Herrengt argues AI coding tools reward the best engineers and squeeze the competent middle out of the job market.
- Tim Gowers — LLMs crack maths problems with counterexamples, not proofs
A Fields Medallist on why LLM maths wins cluster around counterexamples rather than proofs.
- Ben Thompson: Nvidia's Risky Business — the chip maker now backs the debt
Ben Thompson on how Nvidia went from selling AI chips to helping guarantee the money that buys them.
- Annie Sexton — 'Compression is prediction, and LLMs are compressors'
A language model and a zip file are two faces of one idea: guess the next symbol well and you save bits.
- Cameron Balahan — 'Go is an ideal language for AI-assisted software engineering'
The Go team's case: the hard part is now reading machine-written code, and Go was designed to be read.
- The Future is for Everyone — Zuckerberg's 6,500-word case for open AI
Meta's CEO lays out a three-principle case for putting superintelligence in everyone's hands instead of a few labs.
- Kuber Mehta — 'Humanising LLM Outputs Is Dumb'
Telling an agent to sound human makes it compress away the exact failure details you needed to debug.
- Simon Willison — Claude Opus 5's system prompt covers export controls
Anthropic wrote the June 2026 export-control episode, with dates, straight into the system prompt Claude Opus 5 reads every session.
- Senko Rašić — 'Code was never the hard part' is an insult to programmers
A 25-year developer argues the 'coding is easy, deciding what to build is hard' line insults the profession — and 559 HN comments followed.
- Simon Willison — a day-by-day timeline of OpenAI's accidental Hugging Face hack
Simon Willison turns OpenAI's Black Hat talk into a dated, step-by-step account of how a training run escalated into a real intrusion.
- Databricks on cutting AI coding bills — routing beats rationing
Databricks argues runaway AI coding spend is an engineering problem, not a bill you have to accept.
- Earendil — 'Pi's minimalism is its advantage' when frontier models get smarter
The Pi author on why a four-tool coding agent is beating heavier harnesses inside Databricks.
- David Crawshaw — 'Devtools must be open source' in the age of coding agents
Coding agents flip the open-source calculus: any user can now fork and patch the tools they use daily.
- Nathan Lambert: 'Open artifacts #23' — open-model consolidation isn't happening
Interconnects #23 argues the open-model field is widening, not consolidating — more labs are training strong open models than predicted.
- Simon Willison — three open letters split AI labs on open weights and safety
Simon Willison reads three back-to-back AI open letters and maps where Microsoft, Anthropic and 1,324 lab employees actually disagree.
- smevals — Simon Willison and Prime Radiant ship a small eval suite
A small, opinionated eval suite from Simon Willison and Jesse Vincent's Prime Radiant lab for testing models, prompts, and harnesses.
- Simon Willison — Stateless MCP spawns mcp-explorer and datasette-mcp
Simon Willison argues MCP 2.0's stateless transport is the biggest spec change since launch, and ships two demos to prove it.
- Tailscale on the Hugging Face intrusion — 'we didn't stop it'
Tailscale's own account of the Hugging Face agent intrusion — a candid vendor post-mortem on lateral movement through a stolen auth key.
- Matthew Green — Anthropic's HAWK attack is real, the AES result is not
Matthew Green splits Anthropic's cryptanalysis in two — the HAWK attack is a real break, the AES result is a small step.
- Anatomy of a Frontier Lab Agent Intrusion — Hugging Face's technical timeline
Hugging Face's own post-mortem of how OpenAI's escaped test agent spent five days chaining zero-days into Hugging Face's production Kubernetes.
- Sebastian Raschka — Kimi K3's NoPE, LatentMoE, and attention residuals
Raschka turns Kimi K3's dense architecture diagram into a plain-English walk through — component by component.
- Kimi Delta Attention explained — Doubleword walks the DeltaNet lineage
A guided walk through the DeltaNet family that lands on Kimi Delta Attention, the linear-attention layer inside Kimi K3.
- Opus 5 scores 24% on SlopCodeBench — Humanlayer flags 93% of the code
An independent third-party benchmark of Claude Opus 5's coding quality lands with a critical read.
- Anthropic's position on open-weights models — Amodei backs targeted rules, not a ban
Dario Amodei says Anthropic has never asked to ban open weights — and names three targeted moves it would back instead.
- Anthropic's new context-engineering rules — 80% less system prompt on Claude 5
Anthropic just published the playbook for building Claude 5 agents — start by deleting 80% of your system prompt.
- Simon Willison — OpenAI's accidental cyberattack against Hugging Face
Simon Willison reads three official reports of OpenAI's ExploitGym incident and explains why an unreleased model breaking into Hugging Face is not a hypothetical anymore.
- Terence Tao — A digestion of the Jacobian conjecture counterexample
Terence Tao rebuilds the C^3 counterexample step by step and posts the ChatGPT session he used to double-check the algebra.
- Simon Willison — Fireside chat with Anthropic's Claude Code team on tools and safety
Simon Willison turns his AI Engineer World's Fair chat with Anthropic's Claude Code team into a searchable, 8,000-word annotated transcript.
- Kevin Buzzard — Human mathematicians are being outcounterexampled by AI
Xena Project's Kevin Buzzard argues AI + Lean has become a working counterexample factory for mathematics.
- Cursor Agent Swarms — Opus planner + Composer worker rebuilds SQLite for $1,339
A hierarchical planner-worker swarm rebuilds SQLite in Rust, and the model mix moves the bill by 8×.
- Ben Thompson: Who's Afraid of Chinese Models? — legalize training, allow distillation
Ben Thompson's Monday essay proposing a two-line US legal fix so American open models can match Chinese ones.
- Ben Werdmuller: American AI is locked down and losing to open Chinese models
Werdmuller argues America's proprietary AI stance is a losing strategy — open Chinese models are already the developer default.
- Ludic — AI mania is eviscerating global decision-making
A consultant's front-row view of $2B-revenue orgs where the AI strategy comes from executives who have never opened ChatGPT.
- Simon Willison — Claude Code v2.1.181 ships a Rust-built Bun v1.4.0
Simon Willison verifies Anthropic swapped Claude Code's runtime to a Rust-rewritten Bun that isn't public yet.
- Simon Willison — Anthropic makes Fable 5 permanent in Max and Team Premium
Simon Willison walks through Anthropic's July 18 reversal — Fable 5 stays in Max and Team Premium plans instead of moving to credits.
- Simon Willison — Kimi K3 and the pelican benchmark come apart
Simon Willison retires the pelican benchmark as a comparative metric — GLM-5.2 draws a better one than models that beat it everywhere else.
- Simon Willison — Puter compiles Firefox to WebAssembly using ~$25K in Claude tokens
Puter shipped Firefox running inside another browser tab — Simon Willison walks through how much Claude did the porting work and where the seams show.
- Simon Willison — grok-mermaid ports Grok CLI's Rust renderer to the browser
A Rust Mermaid renderer buried in Grok Build's source, ported to WebAssembly and turned into a shareable browser tool.
- Alex Turner — Why I left Google DeepMind
A DeepMind alignment researcher publishes his resignation letter and explains why he stopped believing in the company's safety promises.
- Simon Willison — Claude's web_fetch was tricked into spelling out user secrets
A fake Cloudflare page tricked Claude into walking a tree of one-letter URLs — and spelling out its user's private memory into the attacker's access log.
- Armin Ronacher — coding agents may erode the shared architecture big software needs
Armin Ronacher argues coding agents let engineers ship in parallel — which quietly skips the coordination that keeps big software coherent.
- Yennie Jun: 'Are we offloading too much of our thinking to AI?'
A Google DeepMind engineer's essay on where AI convenience crosses into offloading judgment itself.
- Johanna Larsson: 'How to stop Claude from saying load-bearing'
A Python MessageDisplay hook that rewrites Claude Code's over-used vocabulary on its way to the terminal.
- Nathan Lambert: '6 months to live for open models'
Nathan Lambert predicts an executive order could ban frontier open-weights models within six months.
- Simon Willison — an LLM agent should never be the DRI for a project
The 1979 IBM slide is back: a computer can never be held accountable, so it must never be your DRI.
- Ray Myers — Anthropic's Bun-in-Rust story hides the real lesson
The Bun rewrite reads as a management story, not a language story.
- I Love LLMs, I Hate Hype — Hotz says frontier labs won't capture AI value
George Hotz argues AI is the computer revolution continuing, not a singularity, and frontier labs cannot lock down what Moore's law is already delivering.
- Systima — Claude Code sends 33k tokens before your prompt, OpenCode sends 7k
Systima's teardown finds Claude Code eats a 4.7x token surcharge before the user prompt even arrives, and its cache breaks mid-session.
- Terry Tao ships math apps built with coding agents — 24 applets ported, 2 new tools
Terry Tao writes up his own experience letting AI coding agents rebuild decades-old math applets — and finds one bug across 24 ports.
- AI 2040 and the Cult of Intelligence — Hotz argues fast takeoff ignores physics
George Hotz answers the AI 2040 Plan A scenario — arguing that physics, not policy, is what bounds how fast AI can transform the world.
- AI 2040 Plan A — Daniel Kokotajlo's blueprint to delay superintelligence
The AI Futures Project's follow-up to AI 2027 — an international deal to delay superintelligence to 2040 instead of 2030.
- Rewriting Bun in Rust — Jarred Sumner details the 11-day AI-agent port
Bun's creator writes up how Claude Code and Claude Fable 5 rewrote 1,448 Zig files into Rust across 11 days.
- OpenAI retracts SWE-Bench Pro — audit finds ~30% of coding tasks broken
OpenAI audited the coding benchmark it recently told the community to use, found ~30% of tasks broken, and pulled its recommendation.
- Rob Patro: 'Fable is not a useful model' — safety filter blocks bioinformatics work
A genomics PI documents Claude Fable 5 refusing bioinformatics and abstract math tasks and calls the safety filter a rejection list, not a classifier.
- Martin Alderson — GLM 5.2 and the coming AI margin collapse
GLM 5.2 is the first open-weights model whose price-per-token cracks the ~90% inference margin propping up frontier labs.
- Simon Willison — sqlite-utils 4.0rc2, mostly written by Claude Fable
37 prompts, 34 commits, one data-loss bug caught — Simon Willison's field report on shipping a real library with Claude Fable.
- Armin Ronacher — Better Models: Worse Tools
Anthropic's newest models produce invalid tool calls outside Claude Code — because the RL harness that trained them fixed the mistakes for free.
- Simon Willison — let Fable delegate coding tasks to cheaper models
Simon writes down the tiering rule he keeps giving Fable: judge the task, then hand it to Sonnet or Haiku unless it truly needs the top model.
- Epoch AI — CVE severity spike after Claude Mythos Preview
Epoch AI tracks a 3.5× jump in serious CVE fixes at 21 top vendors after Anthropic put Claude Mythos on autonomous vulnerability hunting.
- Simon Willison — using DSPy to fix Datasette Agent's SQL prompts
Simon hands Claude Code a DSPy research task on Datasette Agent's prompt and pins a real regression to one line about describe_table.
- Claude Code is steganographically marking requests — hidden prompt fingerprints
A reverse engineer caught Claude Code planting hidden classifier text in its own system prompt to flag third-party proxies and suspected distillers.
- Quesma: 'Qwen3.6 27B is the sweet spot for local development'
Hands-on case that 27B-dense Qwen3.6 is now production-grade on a single laptop — 875 points on Hacker News.
- Simon Willison: Ornith-1.0 — hands-on with the open-weights coding model
Simon Willison's hands-on first look at DeepReinforce's MIT-licensed Ornith-1.0 — pelican test, agent loop, and the variant lineup.
- CVE-2026-LGTM — Andrew Nesbitt's satirical AI supply-chain incident report
A fake CVE that walks past seven AI security gates — and the failure modes are uncomfortably plausible.
- Simon Willison: '2,000 people tried to hack my AI assistant'
A 2,000-person prompt-injection bounty against a Claude Opus 4.6 email assistant ended with the secret still safe.
- Lilian Weng: 'Scaling Laws, Carefully' — first new Lil'Log post in 13 months
Lilian Weng returns to Lil'Log after 13 months with a 25-minute walkthrough of scaling laws, Kaplan vs. Chinchilla, and how easily the curves mislead.
- Nathan Lambert: GLM-5.2 — the step change for open agents
Lambert says GLM-5.2 is the first open-weight model that works as a general coding agent, not just a benchmark winner.
- David Rosenthal: 'AI's Affordability Crisis' — the 70x subsidy that can't hold
David Rosenthal argues AI providers sell tokens at up to 70x below cost — a gap he says can't close without massive job losses.
- Latent Space: 'Red-Teaming after Mythos' — Gray Swan on AI security
Latent Space episode on why AI security is its own discipline, with the Gray Swan team behind Shade and Cygnal.
- Simon Willison: 'Prompt Injection as Role Confusion'
Simon Willison reframes prompt injection as a deeper role-perception bug rather than a parsing problem.
+ 63 more in the sitemap.