AI/TLDR — every new AI model, tool, repo & paper
The latest AI releases, refreshed every 2 hours and explained in plain English.
What AI shipped today?
In the last 24 hours AI/TLDR tracked 17 new AI releases, including Suno v6 — a music model family trained only on licensed catalogues, Gander — an open 9B model that listens, watches and works at once and Claude Code 2.1.267 — one setting caps effort on every provider. AI/TLDR is an AI release tracker that follows new AI models, open-source tools, papers, datasets and benchmarks — refreshed every 2 hours from verified primary sources and explained in plain English.
AI Release Index — live stats on AI releases · Learn AI
- Suno v6 — a music model family trained only on licensed catalogues
Suno v6 is a new family of music models built with data licensed from Warner Music Group, BMG and Believe. It ships as v6, v6-wild and v6-mini, and Suno says it will retire its older models as the rollout finishes.
- Gander — an open 9B model that listens, watches and works at once
Gander is an open 9B omni-interaction model that takes streaming video, speech and text together. You can interrupt it mid-sentence, and a separate reasoning agent keeps working on long tasks in the background.
- Claude Code 2.1.267 — one setting caps effort on every provider
Claude Code 2.1.267 adds maxEffortLevel, a setting that caps the effort level on every provider including Bedrock, Vertex and Foundry. Users can still choose a lower level. A long run of fixes repairs prompt-cache reuse.
- Desert Ant Labs — 18 small AI models that run offline inside your app
Desert Ant Labs launched 18 on-device models with Swift, Kotlin and JavaScript SDKs. Each does one job — transcription, PII masking, language detection, video clipping — and runs fully offline. Free up to 100,000 monthly active devices.
- Sebastian Raschka — looped transformers and what Astra's short reasoning means
Sebastian Raschka explains weight-shared 'looped' transformers and pushes back on the idea that GPT-6 Astra hides its reasoning. He argues short chains of thought track model capability, the way bigger models always needed fewer tokens.
- Fireship — 'I built the same game with Astra and Fable 5.1'
Fireship gives GPT-6 Astra and Claude Fable 5.1 the same brief — build a game — and compares what each one produced. The video went up on 9 September 2026, now that Astra is generally available.
- DeepSeek V4.1 Flash — a two-day beta of a natively multimodal rebuild
DeepSeek V4.1 Flash opened as a limited beta on September 8. DeepSeek describes it as an interim model on a new architecture with multimodal input built in, priced like deepseek-v4-flash and capped at 20 concurrent requests per account.
- GPT-5.6 Sol calibrates qubits — Codex agents run measurements at MIT
OpenAI published a case study in which GPT-5.6 Sol, driven through Codex and wired into lab software, ran the standard calibration sequence on an uncalibrated six-qubit superconducting chip in MIT's Engineering Quantum Systems Group.
- TeamAI CLI — Tencent ships one AI setup for a whole team
TeamAI CLI is Tencent's MIT-licensed command-line tool that keeps a team's skills, rules, docs and MCP servers in one git repo and syncs them into Claude Code, Codex, Cursor, CodeBuddy, OpenCode and other agents. It passed 2,700 stars on September 9.
- 1littlecoder — 'open source AI is winning !'
1littlecoder argues in a September 8 video that open-source AI is winning. The week behind it fits the claim: MiniCPM5-2B, IFM's K2 Horizon family and Z.ai's GLM-5.3-Flash all sit in Hugging Face's trending models list.
- Miles v0.1 — RadixArk publishes the technical report for its open RL stack
Miles is an Apache-2.0 reinforcement learning framework for post-training large language and vision models. RadixArk published the v0.1 technical report on September 8, 2026. The repo has 2,685 stars and trains on NVIDIA and AMD accelerators.
- AuK — Tencent's open speech model generates and edits audio by instruction
AuK is an MIT-licensed 1.5B speech model from Tencent Hunyuan, Shanghai Jiao Tong University and the Shanghai Innovation Institute. It generates and edits speech from plain-language instructions. Code and weights shipped on September 8, 2026.
- NeoHorse-1 — open 4B and 9B models post-trained by a routing harness
NeoHorse-1 is a pair of Apache-2.0 models, 4B and 9B, that TokenRhythm post-trained by routing agent work across a pool of models and turning the results into training data. The 4B scores 64.87 on the team's ten-benchmark average, up from 58.94.
- Wes Roth — 'OpenAI JUST solved math....'
Wes Roth's September 9 video works through OpenAI's claim that 10,000 agents produced a proposed solution to the Navier-Stokes Millennium Prize Problem in 88 hours, and the argument over credit that followed it.
- Mercury 2.5 — Inception's diffusion model hits 1,107 tokens per second
Mercury 2.5 is Inception's new diffusion language model, which the company calls the largest ever trained. It runs at 1,107 tokens per second on NVIDIA GPUs, handles 260K tokens of context, and costs $0.20 per million input tokens.
- Deltafin — Kimi K3's 2.8T weights stream off four SSDs at 1 token/s
Deltafin runs the full 2.8-trillion-parameter Kimi K3 on one MacBook Pro. Argonaut Labs' ARGODRIVE build streams 1.45 TB of expert weights from four SSDs into 128 GB of RAM and decodes about 1 token per second.
- Terence Tao — good open math problems are a resource AI is mining out
Terence Tao argues that good open mathematical problems are scarce and slow to replace, and that pointing AI at them indiscriminately solves today's questions while draining the supply that guides the next wave of research.
- Navier–Stokes — OpenAI agents produce a Lean-checked blowup proof
OpenAI says an internal model, running as about 10,000 coordinated agents, produced a proof that 3D Navier–Stokes fluid flow can blow up in finite time. The proof ships as a paper plus a Lean formalization on GitHub.
- Muse — Meta's personal AI agent runs in its own secure virtual machine
Meta launched Muse, a personal AI agent that shops, books travel and fills in forms for you. Each person's Muse runs in a dedicated cloud VM where a separate Sentinel agent has to approve anything that leaves the machine.
- ChatGPT Images 2.5 — OpenAI's image model adds Sketch and two API tiers
ChatGPT Images 2.5 is OpenAI's new image model, with sharper detail, better likeness from reference photos and up to 50% lower generation latency than Images 2.0. Two API models ship with it: GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst.
- Cohere megakernel — one CUDA file serves North Mini Code faster than vLLM
Cohere open-sourced a decode megakernel for North Mini Code that runs 1.25x to 1.41x faster end to end than vLLM on a single H100. It fuses the whole decode step into one persistent CUDA kernel with no compiler dependencies.
- AlphaGenome Atlas — a prediction for every possible DNA letter change
AlphaGenome Atlas is a 1-petabyte dataset from Google DeepMind holding predicted molecular effects for all 9 billion single-letter DNA changes in the human genome. Free to use for non-commercial research from today.
- Two Minute Papers — 'GPT-6 Astra Changes Everything'
Two Minute Papers' September 8 episode works through the first wave of GPT-6 Astra demos that developers posted on X, and pairs them with a new viscous-liquid simulation paper on solving pressure and viscosity together.
- Dan Luu — telling a coding agent to use a test technique barely helps
Dan Luu ran 26 testing instructions and 4 agent skills against the same Rust Zstd task, 80 runs each with Codex on GPT-5.6 Sol. Giving no special instruction at all scored well above average.
- MCP Python SDK 2.2.0 — idle sessions now close after 30 minutes
MCP Python SDK 2.2.0 hardens Streamable HTTP: idle sessions close on their own, a server holds at most 10,000 at once, and HTTP redirects only follow within the same origin. The same changes shipped as 1.30.0 on the 1.x line.
- UltraData-RL-2609 — 86,000 checkable RL tasks behind MiniCPM5-2B
UltraData-RL-2609 is an Apache-2.0 reinforcement-learning corpus of 85,995 tasks in math, code, long-context and knowledge. Every item carries a reference answer and a defined way to check it. OpenBMB used it to post-train MiniCPM5-2B.
- Fireship — 'Big AI wants you broke... here are some free alternatives'
Fireship runs through five free, open-source tools for cutting the token bill on AI coding work: Ollama, 9router, Headroom, Diffy and OpenHands. The video went up on 7 September 2026.
- MiniCPM5-2B — a 2B open model that leads the sub-4B field
MiniCPM5-2B is OpenBMB's new 2.52B-parameter open model for phones and laptops. It averages 53.9 across the maker's benchmark set, ahead of Qwen3.5-4B at 51.1, and ships under Apache-2.0 with a 131K-token context.
- Wes Roth — 'OpenAI's chief scientist just issued a warning'
Wes Roth works through An Alien Mind, the September 6 essay by OpenAI chief scientist Jakub Pachocki, and asks whether alignment work can keep pace with recursive self-improvement. He also covers OpenAI's research-acceleration data.
- Humanizer v3.0.0 — the AI-writing cleanup skill drops 35 patterns to 25
Humanizer v3.0.0 rewrites AI-sounding text as an agent skill for Claude Code and Claude Desktop. The release merges 35 writing patterns into 25 without dropping any checks, and adds a validator that caps the skill file at 400 lines.
- OpenMAIC 1.0.1 — four security advisories, one rated critical
OpenMAIC 1.0.1 fixes four privately reported security holes in the open-source multi-agent classroom, including a critical unauthenticated SSRF that could reach a cloud metadata service. Anyone running 1.0.0 should upgrade.
- LiteLLM v1.100.0 — Vertex AI Interactions API and shared team budgets
LiteLLM v1.100.0 adds native Vertex AI Interactions API support, Bing Search grounding, day-0 Gemini transcription models, and shared budgets across model access groups. It references 398 pull requests and deletes prompt_token_calculator.
- Bryan Cantrill — readers quit a post the moment they spot AI writing
Bryan Cantrill argues that LLM-written prose repels the readers a writer most wants. He cites a survey of 668 developers by Cynthia Dunlop in which 78% stop reading as soon as they notice AI writing and 71% avoid that author afterwards.
- Ollama 0.34 — local models now run inside ChatGPT Desktop
Ollama 0.34, published on September 5, 2026 as a release candidate, lets ChatGPT Desktop answer using models that run on your own machine. Setup happens in the Ollama macOS app. The release also speeds up structured output on Apple Silicon.
- An Alien Mind — OpenAI's chief scientist calls for voluntary slowdowns
Jakub Pachocki, OpenAI's chief scientist, published an essay saying no lab has solved alignment and monitoring well enough to keep scaling at full speed. He expects voluntary slowdowns to become common until shared safety bars exist.
- Research acceleration at OpenAI — 3.1 agent workdays per human workday
OpenAI published numbers on how coding agents changed its own research. By mid-August the research org used 3.1 agent-workdays for every human workday, and OpenAI says it reached its goal of an automated research intern by September 2026.
- Lyria 3.5 comes to Gemini — Google's music model lands in the app and API
Lyria 3.5 is Google's music generation model, and it is now in the Gemini app for every user worldwide as well as in the Gemini API. Developers call it as the model id lyria-3.5 and get 44.1 kHz stereo tracks of up to about three minutes.
- Sam Witteveen — 'NVIDIA Doubles Down on Local AI With PAIR'
Sam Witteveen walks through NVIDIA PAIR, the Personal AI Router NVIDIA announced at IFA 2026 on 3 September. The Apache-2.0 tool finds the other PCs on your network and sends each local inference request to whichever machine has a free GPU.
- OpenClaw 2026.9.2 — GPT-6 Astra support and swarms on by default
OpenClaw 2026.9.2 adds GPT-6 Astra and Muse Spark 1.3, turns sub-agent swarms on by default, and applies most settings without a Gateway restart. Session tools now let agents read each other's conversations unless you narrow them.
- GPT-6 Astra on robot arms — 95% on a block task, Claude Fable 5.1 got 40%
Robocurve, an independent robotics benchmark group, gave GPT-6 Astra and Claude Fable 5.1 control of the same robot arms. Astra put the red block in the bowl 19 times out of 20; Fable 5.1 managed 8, at over twice the cost per run.