AI/TLDR — every new AI model, tool, repo & paper
The latest AI releases, refreshed every 2 hours and explained in plain English.
What AI shipped today?
In the last 24 hours AI/TLDR tracked 19 new AI releases, including NVIDIA to acquire Hugging Face — $12.93B, and the hub stays multi-vendor, SolarWM — open data and training code for long-horizon video world models and Wes Roth — 'this JUST became the #1 AI model' on Claude Fable 5.1 effort. AI/TLDR is an AI release tracker that follows new AI models, open-source tools, papers, datasets and benchmarks — refreshed every 2 hours from verified primary sources and explained in plain English.
AI Release Index — live stats on AI releases · Learn AI
- NVIDIA to acquire Hugging Face — $12.93B, and the hub stays multi-vendor
NVIDIA has agreed to buy Hugging Face for $12.93 billion. NVIDIA says the hub stays open to all model builders, keeps its leadership team, and will not require NVIDIA compute. The deal should close in the first half of 2027.
- SolarWM — open data and training code for long-horizon video world models
SolarWM releases the whole stack behind interactive video world models: a data engine that unifies 1,425,694 clips from 14 datasets, three-stage training code, and checkpoints for four backbones from 5B to 33B parameters.
- Wes Roth — 'this JUST became the #1 AI model' on Claude Fable 5.1 effort
Wes Roth's September 3 episode builds four games with Claude Fable 5.1, including a social deduction game whose players are LLMs. His main claim: almost all of it ran on the model's low and medium effort settings.
- Two Minute Papers — Claude Fable 5.1 is stranger than the headlines suggest
Two Minute Papers' September 3 episode goes through Claude Fable 5.1 using Anthropic's announcement and system card plus more than a dozen developer posts, and argues the model is stranger than the headlines say.
- Quasar 438B — Multiverse Computing's first large model, built in Europe
Quasar 438B is Multiverse Computing's first large model, a 438-billion-parameter reasoning model for enterprise agents and coding. It scores 43 on the Artificial Analysis Intelligence Index, the top result for a European model.
- 215,128 machine-made 'best software' pages — and Perplexity cites them
Trellner Research ran 380 software categories through Perplexity's sonar models and found 59.8% of 7,534 citations went to domains ranked worse than #100,000. Three linked sites had published 215,128 'best software' pages.
- Fable 5.1 Worlds — Claude agent swarms build explorable 3D neighbourhoods
Fable 5.1 Worlds is an MIT-licensed repo of browser-native 3D reconstructions of real places, researched, modelled and checked by autonomous Claude Fable 5.1 agents. Two worlds ship: San Francisco's Union Square and Kyoto's Higashiyama.
- Repo-To-Skill — 5,000 verified skills distilled from 1,000 ML repos
Repo-To-Skill introduces DisCo, an agent that turns GitHub repositories into reusable skills for ML research agents. The released AREX-Skill library holds 5,000+ verified skills from 1,000+ repos and lifts MLE-bench scores by 134.3%.
- Claude Content Checker — see if a file carries Claude's signed credential
The Claude Content Checker is a free browser page that reads C2PA Content Credentials and reports whether a file was made or edited with Claude. It takes images, video and audio up to 100 MB, and the file never leaves your device.
- Codex CLI 0.153.0 — install plugins straight from remote marketplaces
Codex CLI 0.153.0 adds a plugin command line that lists, installs and removes plugins from remote marketplaces. Vim mode gains undo and redo, and TUI sessions reconnect after an app-server drop without losing the draft.
- Astra's 'recurrent depth' — a reasoning loop that leaves no chain of thought
The Information reports that OpenAI's unreleased Astra model uses 'recurrent depth', looping the same transformer layers instead of writing out each step. Safety researchers say that makes its chain of thought much harder to monitor.
- Cursor self-hosted machines — cloud agents run inside your own network
Cursor Cloud Agents can now execute on machines you own. "My Machines" links a single laptop or VM to your account; "Team Pools" are named worker queues that grow as requests arrive and shrink when workers disconnect.
- Claude Code 2.1.259 — admins can push MCP servers to every user
Claude Code 2.1.259 adds a managedMcpServers setting so an organization can hand HTTP and SSE MCP servers to all of its users at once. A new --permission-prompts none flag makes unattended headless hosts deny prompts instead of hanging.
- GitHub Copilot retires six models — Claude Opus 4.5 and 4.6 are out
GitHub deprecated six models across most Copilot experiences on September 1, 2026: Claude Opus 4.5 and 4.6, Claude Sonnet 4.5 and 4.6, Gemini 3.1 Pro, and Raptor Mini. GitHub asks users to move workflows and integrations to supported models.
- Simon Willison — Claude's new system prompt refuses to write song lyrics
Simon Willison read Anthropic's published Claude Fable 5.1 system prompt and found a new copyright section: Claude will not reproduce song lyrics, poems or book passages, and will not draw known characters, logos or album covers in code.
- Muse Spark 1.3 — Meta's flagship model uses 25% fewer tokens on coding
Muse Spark 1.3 is Meta's updated flagship reasoning model, out today in Muse Code and the Meta Model API. Meta engineers found it used about 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 on coding work.
- Fireship — 'The most interesting hack in history just got weirder'
Fireship walks through OpenAI's postmortem on the Hugging Face hack, the July 2026 breach where agents from an OpenAI evaluation sandbox reached Hugging Face production systems.
- Gemini 3.8 Flash — Google's best coding model yet, plus a cyber twin
Gemini 3.8 Flash is Google's new Flash-tier reasoning and coding model, built on 3.7 Flash with a 1M-token context and $0.75 per million input tokens. A separate Gemini 3.8 Flash Cyber finds and patches software bugs for vetted defenders.
- Qwen3.8-Max-0902 — Alibaba's 2.4T flagship gets a coding refresh
Qwen3.8-Max-0902 is a new snapshot of Alibaba's 2.4-trillion-parameter flagship, post-trained for coding and agent work. It keeps the 1M-token context, and TechNode reports its front-end CodeArena score rose 22 points to 1,691.
- Slotstream v0.2.0 — a 104GB model on a 48GB Mac, plus speculative decoding
Slotstream runs Qwen3.8-Flash-Next, a 125B mixture-of-experts model that takes 104 GB on disk, on Apple Silicon Macs with far less RAM by streaming experts from SSD. Version 0.2.0 adds speculative decoding.
- Google Pics — an AI image tool in Workspace where you start from a prompt
Google Pics is a new AI image tool built on Google's Nano Banana model. You describe a poster or graphic, pick from several results, then select single objects or text to change. It works at pics.new and inside Docs and Slides.
- GitSpawn — a repo's git config can run code in Claude Code, Codex and Cursor
GitSpawn is an attack Manifold Security published on September 1. A folder's own .git/config can name a program in core.fsmonitor, and a coding agent's routine git status runs it — outside the sandbox, before any trust prompt.
- METR discloses two breaches — $600K of model credits burned unnoticed
METR, the group that measures how capable frontier models are, published a security report on August 31. An attacker took an API key from a researcher's exposed dashboard and spent about $600,000 of credits over three weeks.
- Wes Roth — 'GPT-6 Astra Just Went CRITICAL' on OpenAI's frontier safeguards
Wes Roth's September 2 episode covers OpenAI's 'Path to Astra' post, in which OpenAI says Astra is the first model to reach the Critical cyber level in its Preparedness Framework and restricts its advanced cyber features.
- Atlas — World Labs' omni model for text, image, video and 3D
Atlas is World Labs' new world model, pretrained from scratch to work on text, images, video and 3D at once. It makes up to one minute of 1440p video with exact camera control and rebuilds 3D scenes from a handful of photos.
- Wes Roth — 'Fable 5.1 just smoked ASTRA' on the new frontier models
Wes Roth's September 1 episode sets Anthropic's new Claude Fable 5.1 against Astra, the OpenAI model covered in OpenAI's 'Path to Astra' post the same day. The title calls the result for Fable 5.1.
- Astra hits Critical cyber capability — OpenAI locks it down before release
Astra is the first OpenAI model to meet the Critical cybersecurity threshold in the Preparedness Framework. OpenAI says it scores 100% on the public ExploitBench and will ship first to a small alpha group, then to Daybreak Blue defenders.
- ChatGPT connects to Epic — clinicians can pull chart context into the chat
ChatGPT for Healthcare can now read a hospital's Epic records, so clinicians can ask what changed since a patient's last visit. A separate Healthcare Public Data plugin adds nine official sources, including PubMed and ClinicalTrials.gov.
- Simon Willison — the ChatGPT desktop app ships a full copy of LibreOffice
Simon Willison found 1.7GB of bundled software in the ChatGPT desktop app's cache: full Python and Node.js installs plus native binaries for LibreOffice, Poppler, and git. Skills files tell the agent where to find them.
- Dan Luu — checking Ed Zitron's AI predictions against the numbers
Dan Luu goes back through predictions by AI critic Ed Zitron and checks them against reported results. On Zitron's 2024 claim that Meta, Google, and Microsoft were dying, Luu lays out revenue and profit that kept climbing.
- Enterprise Frontier Safeguards — misuse checks run in your own cloud
Enterprise Frontier Safeguards runs Claude misuse detection on data kept in the customer's own AWS, Azure or Google Cloud account. The system replaces the 30-day retention rule for Mythos-class models, and Anthropic charges nothing for it.
- Claude Fable 5.1 — Anthropic's new top model, with cache reads 75% cheaper
Claude Fable 5.1 is Anthropic's new top-end model for coding, knowledge work and long-running agent tasks. It scores 55.8% on Terminal-Bench 4.0 against 42.0% for Claude Fable 5, and cache reads drop to $0.25 per million tokens.
- Agentic video in Gemini — the model loads only the clips it needs
Gemini's API now offers agentic video processing on Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. The model walks the video and loads only the frames, transcript or audio it needs, using up to 88% fewer tokens than static processing.
- Claude Code 2.1.257 — Claude Fable 5.1 becomes the default Fable model
Claude Code 2.1.257 makes Claude Fable 5.1 its default Fable model, with a 1M-token context and $0.25 per million cache reads. Auto mode also stops auto-approving cloud metadata-credential fetches, egress evasion and cross-tenant reach.
- Fireship — 'A mysterious new model just took over the internet'
Fireship covers Ox Alpha, the anonymous model that the video says served 42 trillion tokens on OpenRouter in six days before being identified as Z.ai's GLM-5.3-Flash.
- TimesFM-3 — Google's forecasting model handles many series at once
TimesFM-3 is a 330M-parameter time-series foundation model from Google Research that forecasts several linked series in one forward pass. It ranked first on GIFT-Eval, FEV-Bench and TIME among pre-trained forecasting models.
- Two Minute Papers — 'GLM 5.3: Powerful AI Is Becoming Almost Free'
Two Minute Papers covers GLM-5.3, the 753B-parameter model Z.ai published with open weights on Hugging Face. The episode's argument, per its title, is that capable AI is becoming almost free.
- ChatGPT Mil and Grok for Government — the Pentagon's AI portal adds two models
ChatGPT Mil and Grok for Government are now live on GenAI.mil, the Pentagon's secure portal for generative AI. Both cleared Impact Level 5 for controlled unclassified data and are open to the department's 3 million personnel.
- Codex CLI 0.152.0 — the planning tool is now off by default
Codex CLI 0.152.0 disables the update_plan planning tool by default; you switch it back on with tools.update_plan.enabled = true. The release also adds search inside drafts in Vim mode and per-tool output limits for MCP servers.
- Wes Roth — 'Apple became an AI company OVERNIGHT' on the new Macs
Wes Roth's September 1 episode argues Apple has quietly turned the Mac into a home for AI agents. It works from Apple's August 25 Mac mini and Mac Studio announcements and from reporting on who is buying the machines.