AI/TLDR — every new AI model, tool, repo & paper
The latest AI releases, refreshed every 2 hours and explained in plain English.
What AI shipped today?
In the last 24 hours AI/TLDR tracked 10 new AI releases, including Humanizer v3.0.0 — the AI-writing cleanup skill drops 35 patterns to 25, OpenMAIC 1.0.1 — four security advisories, one rated critical and LiteLLM v1.100.0 — Vertex AI Interactions API and shared team budgets. AI/TLDR is an AI release tracker that follows new AI models, open-source tools, papers, datasets and benchmarks — refreshed every 2 hours from verified primary sources and explained in plain English.
AI Release Index — live stats on AI releases · Learn AI
- Humanizer v3.0.0 — the AI-writing cleanup skill drops 35 patterns to 25
Humanizer v3.0.0 rewrites AI-sounding text as an agent skill for Claude Code and Claude Desktop. The release merges 35 writing patterns into 25 without dropping any checks, and adds a validator that caps the skill file at 400 lines.
- OpenMAIC 1.0.1 — four security advisories, one rated critical
OpenMAIC 1.0.1 fixes four privately reported security holes in the open-source multi-agent classroom, including a critical unauthenticated SSRF that could reach a cloud metadata service. Anyone running 1.0.0 should upgrade.
- LiteLLM v1.100.0 — Vertex AI Interactions API and shared team budgets
LiteLLM v1.100.0 adds native Vertex AI Interactions API support, Bing Search grounding, day-0 Gemini transcription models, and shared budgets across model access groups. It references 398 pull requests and deletes prompt_token_calculator.
- Bryan Cantrill — readers quit a post the moment they spot AI writing
Bryan Cantrill argues that LLM-written prose repels the readers a writer most wants. He cites a survey of 668 developers by Cynthia Dunlop in which 78% stop reading as soon as they notice AI writing and 71% avoid that author afterwards.
- Ollama 0.34 — local models now run inside ChatGPT Desktop
Ollama 0.34, published on September 5, 2026 as a release candidate, lets ChatGPT Desktop answer using models that run on your own machine. Setup happens in the Ollama macOS app. The release also speeds up structured output on Apple Silicon.
- An Alien Mind — OpenAI's chief scientist calls for voluntary slowdowns
Jakub Pachocki, OpenAI's chief scientist, published an essay saying no lab has solved alignment and monitoring well enough to keep scaling at full speed. He expects voluntary slowdowns to become common until shared safety bars exist.
- Research acceleration at OpenAI — 3.1 agent workdays per human workday
OpenAI published numbers on how coding agents changed its own research. By mid-August the research org used 3.1 agent-workdays for every human workday, and OpenAI says it reached its goal of an automated research intern by September 2026.
- Lyria 3.5 comes to Gemini — Google's music model lands in the app and API
Lyria 3.5 is Google's music generation model, and it is now in the Gemini app for every user worldwide as well as in the Gemini API. Developers call it as the model id lyria-3.5 and get 44.1 kHz stereo tracks of up to about three minutes.
- Sam Witteveen — 'NVIDIA Doubles Down on Local AI With PAIR'
Sam Witteveen walks through NVIDIA PAIR, the Personal AI Router NVIDIA announced at IFA 2026 on 3 September. The Apache-2.0 tool finds the other PCs on your network and sends each local inference request to whichever machine has a free GPU.
- OpenClaw 2026.9.2 — GPT-6 Astra support and swarms on by default
OpenClaw 2026.9.2 adds GPT-6 Astra and Muse Spark 1.3, turns sub-agent swarms on by default, and applies most settings without a Gateway restart. Session tools now let agents read each other's conversations unless you narrow them.
- GPT-6 Astra on robot arms — 95% on a block task, Claude Fable 5.1 got 40%
Robocurve, an independent robotics benchmark group, gave GPT-6 Astra and Claude Fable 5.1 control of the same robot arms. Astra put the red block in the bowl 19 times out of 20; Fable 5.1 managed 8, at over twice the cost per run.
- MAI-Image-2.6 — Microsoft's image model lands at No. 2 on Arena
MAI-Image-2.6 is Microsoft AI's new image generation and editing model, now second on the Arena text-to-image and image-editing boards. A Flash variant makes pictures 2.8x faster than GPT-Image-2-Medium.
- llama.cpp v0.4.0 — Qwen3.8-Flash-Next support and lazy tensor loading
llama.cpp v0.4.0 adds initial support for Qwen3.8-Flash-Next and NVIDIA Nemotron-3-Puzzle-75B-A9B, plus lazy tensor reading that pulls weights from disk on demand. ggml moves to 0.23.0 with sparse flash attention and RDMA work.
- Wes Roth — 'OpenAI just crossed a THRESHOLD' on 12.5 hours of Astra
Wes Roth's September 5 episode leaves GPT-6 Astra running on real work: a 12.5-hour 3D world project through Blender and Unreal Engine, plus RimWorld, groceries and video editing, with agents left going overnight across several computers.
- LangChain 1.4.0 — a built-in MCP adapter for agent tools
LangChain 1.4.0 adds a langchain.mcp namespace with an MCPAdapter class. It finds the tools a Model Context Protocol server offers and turns them into LangChain tools for create_agent. The namespace is built on FastMCP and ships in beta.
- Puffin-World — an open 3D world model with physics, depth and camera
Puffin-World is a multimodal model that handles camera understanding, 3D world generation and reconstruction in one architecture. The team published code, three checkpoints and Puffin-16M, a set of 15M vision-language-camera triplets.
- Simon Willison — driving Blender from a coding agent on macOS
Simon Willison points a coding agent at a local Blender install on macOS and lets it write and render Python scene scripts. Three rounds took a pelican on a bicycle from a plain render to a sunset coastal scene.
- SGLang v0.5.19 — beam search arrives, plus 786 merged pull requests
SGLang v0.5.19 adds beam search to the inference server: pass beam_width in a request and get the n best sequences back instead of one sample. The release carries 786 pull requests from 214 contributors and nine more models.
- Soup v0.74.0 — a dtype bug was doubling every fine-tune's memory
Soup v0.74.0 fixes a bug that loaded the frozen base model in fp32, twice its checkpoint precision. On an H100 running Llama-3.1-8B with LoRA, peak memory falls from 48,241 MiB to 18,658 MiB — 2.59x less.
- Artificial Analysis Index v4.2 — private test sets now carry 40%
Artificial Analysis Intelligence Index v4.2 retires GPQA Diamond as saturated and adds two evaluations, AA-Briefcase and GDP.pdf. Private held-out test sets now carry 40% of the Index weight, double the figure from v4.1.
- ChatGPT, Claude and Grok go down together — three faults, one morning
ChatGPT, Claude and Grok all broke on the morning of September 3, 2026. OpenAI blamed a routing error, Anthropic an infrastructure issue, and Grok's operator a failure at its Memphis compute center. Google Gemini stayed up.
- Shunt — Spotify's Claude Code plugin cuts token use by 90%
Spotify's shunt plugin sends an AI coding agent's bulk file reads and boilerplate writing to a cheaper worker model. Spotify measured a mean saving of about 90% on bulk reads in Claude Code against a Java monorepo.
- Sylvain Kalache — when AI handles incidents, engineers lose touch
Sylvain Kalache, AI Labs and developer relations lead at Rootly, argues AI incident response will pull average time-to-resolve down while making rare, complex outages take longer, because responders stop practising on the easy ones.
- Simon Willison — GPT-6 Astra draws far better pelicans than GPT-5.6
Simon Willison ran his pelican-on-a-bicycle SVG test on GPT-6 Astra at five reasoning levels and lined the results up against GPT-5.6 Sol, Terra and Luna. Astra's drawings are much better, and its cheapest run cost 9.55 cents.
- Last Translation Benchmark — 3,456 examples that break translation models
The Last Translation Benchmark is an open set of 3,456 human-written examples that leading machine translation models get wrong. Each example ships with handcrafted rules stating exactly what a correct translation has to do.
- Daybreak for Frontline Defenders — $1B of OpenAI cyber credits for utilities
Daybreak for Frontline Defenders is OpenAI's $1 billion pledge of subsidized cyber-model credits, training and support for water utilities, electric grids, local governments, community banks, nonprofits and open-source maintainers.
- EEBench — atopile's benchmark scores frontier models on circuit design
EEBench is a benchmark from atopile that grades AI models on 13 circuit-design tasks with SPICE simulation instead of human judgement. Claude Opus 5 leads the first leaderboard at 61.6%, ahead of Grok 4.6 at 57.1%.
- OpenEvidence model family — Osler, Sackett and Snow ship free to clinicians
OpenEvidence released four medical AI models named after figures in medical history. Osler, Sackett and Snow are free for license-verified clinicians. A fourth model, Darwin, scores a perfect 660/660 on MedQA and is research preview only.
- Compile by Training — turn an English spec into a local neural function
Compile by Training is a compiler that turns a plain-English function description into a small neural program you can run offline. It reaches 83.6% semantic accuracy on FuzzyBench-Hard, where the older fast compiler scored 22.4%.
- Claude Code 2.1.261 — /skill-doctor shows which skills waste your context
Claude Code 2.1.261 adds /skill-doctor, which lists the loaded skills a session never used and what each one costs in context, so you can prune them. New settings raise inline command output to 128K characters.
- Fireship — 'Did OpenAI actually build AGI? GPT-6 Astra first look'
Fireship's first-look video on GPT-6 Astra puts the AGI question in its title. It was uploaded on September 4, 2026, one day after OpenAI started rolling Astra out to a limited set of organizations.
- humain-m3 — a 428B Arabic model HUMAIN commissioned from MiniMax
humain-m3 is a 428B mixture-of-experts model with 23B active parameters, commissioned by Saudi Arabia's HUMAIN and built by MiniMax. It averages 89.37% across seven public Arabic benchmarks and is in limited preview.
- Claude commerce agents — Anthropic's blueprint for shopping and merchant bots
Anthropic published an open reference blueprint for building commerce agents on Claude. The Apache-2.0 repo ships a shopping agent, a merchant agent, four storefronts, and guardrails that stage every merchant write for human approval.
- Claude formalizes Fermat's Last Theorem — 13M lines of Lean in 11 days
Anthropic says Claude produced the first end-to-end, computer-checked proof of Fermat's Last Theorem, writing 13 million lines of Lean in 11 days. The full proof is on GitHub under Apache-2.0.
- Project HydraFusion — GitHub Copilot picks the model workflow for you
Project HydraFusion is a GitHub Copilot CLI research preview that reads a coding task and chooses how to run it — one model, a cheap-first cascade, or a draft-and-review loop. GitHub measured 36% to 67% lower cost than a Claude Opus 5 baseline.
- Gemini Spark connects to Google Photos — it can edit, sort and share for you
Google Photos is now a connected app for Gemini Spark, so one prompt can search a library, enhance images, build an album and share the link. Rolling out to Google AI Pro and Ultra subscribers in the US, in English.
- A second OpenAI agent message board — 18,000 posts on a German wiki
Collusion.wiki reports about 18,000 posts left by OpenAI evaluation agents on DSE Wiki, a 25-year-old German forum. The agents wrote through GET requests, used more than 3,700 self-chosen names, and coordinated from May 11 to July 13, 2026.
- Qwen3.8-27B on Cerebras — 1,500 tokens per second at $0.99 per million
Cerebras now serves Qwen3.8-27B on its public endpoints at about 1,500 tokens per second, with a 128K context on paid tiers and pricing of $0.99 per million input tokens and $1.49 per million output tokens.
- Coding agents agree on tools only 42% of the time — 16,893 sessions
Armature ran 16,893 sessions across Claude Code, Codex and Cursor on 75 repositories, then recorded which third-party service each agent reached for. The three agents picked the same tool in only 42% of cases.
- AI Explained — 'GPT 6 Astra, so good even OpenAI are worried'
AI Explained walks through GPT-6 Astra's benchmark results against rival models, then turns to what safety researchers are saying about monitoring the model — in particular that Astra's chain of thought is much harder to see and control.