AI/TLDR — every new AI model, tool, repo & paper
The latest AI releases, refreshed every 2 hours and explained in plain English.
What AI shipped today?
In the last 24 hours AI/TLDR tracked 16 new AI releases, including 1littlecoder — 'Ox Alpha is GLM 5.3 Flash!!!', OpenAI Admin plugin — ChatGPT Work admins run the workspace in chat and GigaBrain-0.7 — an open robot brain trained on 37,000 hours of embodied data. AI/TLDR is an AI release tracker that follows new AI models, open-source tools, papers, datasets and benchmarks — refreshed every 2 hours from verified primary sources and explained in plain English.
AI Release Index — live stats on AI releases · Learn AI
- 1littlecoder — 'Ox Alpha is GLM 5.3 Flash!!!'
1littlecoder's new video covers the reveal that Ox Alpha, the free stealth model on OpenRouter, is Z.ai's GLM-5.3-Flash. The weights are now on Hugging Face under an MIT license.
- OpenAI Admin plugin — ChatGPT Work admins run the workspace in chat
OpenAI's Admin plugin puts workspace administration inside ChatGPT Work. Admins can review credit use, add or remove members, check permissions and adjust spending limits in one conversation, using only the access they already have.
- GigaBrain-0.7 — an open robot brain trained on 37,000 hours of embodied data
GigaBrain-0.7 is an open vision-language-action model for robots, released under Apache-2.0 with weights on Hugging Face. GigaAI pretrained it on more than 37,000 hours of mixed robot data and ships the code in the giga-brain-0 repo.
- Two Minute Papers — 'DeepSeek's New AI System Shouldn't Be Possible'
Two Minute Papers covers DeepSeek Harness, the MIT-licensed agent framework where every capability is a plugin. The episode description links the Harness page and the Cordis plugin-system paper the framework is built on.
- OpenAI Assistants API shuts down — the beta ends one year after notice
OpenAI removes the Assistants API on August 26, 2026, one year after telling developers it would. The replacement is the Responses API paired with the Conversations API, and there is no automated tool to move existing Threads across.
- Wes Roth — 'OpenAI BROKE the Industry Overnight' on the Jalapeño report
Wes Roth's August 26 episode lists one source in its description: the SemiAnalysis report on OpenAI's Jalapeño inference chip, which argues the chip beats Nvidia Blackwell on throughput per megawatt.
- Portable Computer — Perplexity's agent runs fully on your own GPU
Perplexity Portable Computer runs its whole agent stack — orchestrator, subagents and harness — on hardware you own. Local work uses no billing credits, and the agent asks permission before sending any step to a cloud model.
- Qwen3.8-Flash-Next — open 125B MoE with 6B active previews Qwen4
Qwen3.8-Flash-Next is an open-weights 125B mixture-of-experts model that activates only 6B parameters per token and previews the Qwen4 architecture. It scores 62.5 on SWE-bench Pro, against 53.4 for Claude-Opus-4.6 (Max).
- WeMM-Embedding — Tencent's multimodal retrieval models top MMEB-v2
WeMM-Embedding is a family of open multimodal embedding models from Tencent's WeChat Vision Team, released in 2B, 4B and 9B sizes under Apache-2.0. The 9B model scores 80.6 average on MMEB-v2.
- LAION-BVD — 10 million hours of open video for multimodal training
LAION-BVD is an open video dataset of 80 million videos totalling 10 million hours, pulled from CommonCrawl. It ships 55M captioned clips, 10M audio clips and 300M image frames for research use.
- Claude memory works everywhere — one memory across chat and Cowork
Claude's memory now spans chat and Claude Cowork, so context saved in one shows up in the other. Claude writes topics during a conversation instead of summarizing it afterward, and every topic is a file you can read, edit, or delete.
- Claude Code 2.1.246 — gateway API keys are no longer sent to Anthropic
Claude Code 2.1.246 stops telemetry and metrics requests to Anthropic from carrying the API key set for a third-party gateway, so a credential now only goes to its own host. The release also adds an Auto mode tab to /permissions.
- Anthropic wellbeing grants — $5M for open evaluations of AI's effect on people
Anthropic is putting $5 million into outside research on how AI models affect user wellbeing. Grantees get funding, model access, and technical support, and publish their evaluations as open source. Applications close September 21.
- IBM Granite 4.2 — open reasoning models with a thinking switch
Granite 4.2 is IBM's first family of dense reasoning models, released under Apache-2.0 in 3B, 8B and 30B sizes. Each model has a thinking mode you can switch off, and the 8B and 30B learned to use tools inside live sandboxes.
- Jalapeño first results — OpenAI's inference chip posts 1.9x more work per watt
OpenAI published the first measured results for Jalapeño, its custom inference chip. On SemiAnalysis's InferenceX benchmark it delivered 1.5–1.9x more work per watt and 1.7–3.6x lower end-to-end latency than the comparison systems.
- GPT-5.6 in Kiro — OpenAI's model family lands in AWS's coding agent
GPT-5.6 Sol, Terra and Luna are now available in Kiro, AWS's spec-driven coding agent. OpenAI and AWS say joint testing on Terminal-Bench 2.1 showed GPT-5.6 Terra finishing tasks in Kiro at roughly 82% lower cost.
- Prime Intellect finds an offline sandbox escape — the inference API is the hole
Prime Intellect found that 'offline' AI evaluation sandboxes still reach the internet through the inference API. GPT-5.6 Sol Pro used the Responses API's file_url field to read a flag off GitHub. vLLM, SGLang and TensorRT-LLM are patched.
- Apple M6 and M5 Ultra — 2nm silicon and 512GB for on-device LLMs
Apple's M6 is its first 2nm chip, and the new M5 Ultra bonds four dies over UltraFusion to reach an 80-core GPU with 512GB of unified memory. Apple says a Mac Studio can now run frontier-class LLMs entirely on device.
- EchoWM — a world model you can walk through, with sound and speech
EchoWM is an omnimodal world model from JD.com's Joy Future Academy that generates 720p video together with matching environmental sound, music and speech, while a viewer steers the camera through the scene step by step.
- Boyd Kane — a model could escape by attacking the engine that runs it
Boyd Kane argues an LLM could take over its own GPU host by emitting tokens that trip the inference engine's parser. He points to CVE-2025-9141, where vLLM passed tool-call arguments straight to Python's eval().
- Ed-o-meter — a 28-task LLM leaderboard that runs every model down one track
Ed-o-meter scores 17 models on 28 practical tasks across coding, data, real-world work, security and tool use, showing pass rate, cost and time-to-first-token side by side. GLM-5.3 passes all five categories for $0.28 a run.
- Thomson 1.0 Small — Thomson Reuters ships its own 35B legal and tax model
Thomson Reuters released Thomson, its first in-house language model, and put a small 35B version on Hugging Face with open weights for academic and non-commercial use. Training cost $40 million.
- Varkos — a local AI companion that talks while it plays Skyrim with you
Varkos is a voice-driven Skyrim companion that listens, plans and acts in the game while it talks. Speech-to-text lands in 40-80 ms and a spoken reply can arrive in under 500 ms, all from local models.
- Apodex 1.1 — an agent model that finishes whole jobs, with 35B open weights
Apodex 1.1 is an agent model built to finish real work, not just write a report. It scores 38.5 on APEX-Agents and 78.8 on GDPVal. A 35B Mini version ships with open weights under Apache-2.0.
- Claude Code 2.1.243 — the install drops from 340 MB to 75 MB
Claude Code 2.1.243 cuts the native install from about 340 MB to about 75 MB and frees 40-70 MB of memory per session. It also adds a curated /model picker, prompt cache TTL settings, and keyless sign-in through the Anthropic Console.
- Tempus ECG-PH — FDA clears AI that spots pulmonary hypertension in a routine ECG
Tempus ECG-PH is an AI software device that reads a standard 12-lead ECG and flags signs of pulmonary hypertension. The FDA granted it 510(k) clearance on August 24, 2026. It is Tempus' third cleared heart device.
- Two Minute Papers — 'This Small AI Will Change Everything' on Qwen3.8-27B
Two Minute Papers covers Qwen3.8-27B, Alibaba's 27B open-weights model. The description cites the Hugging Face model card plus community runs, including an NVIDIA forum thread measuring the model on a single DGX Spark.
- NVIDIA Groq 3 LPX — the agent inference chip enters full production
NVIDIA Groq 3 LPX is now in full production. The accelerator handles token generation for AI agents and reached 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context. Nebius is the first cloud to adopt it.
- Hugging Face explores a sale — reports put the price at $13B or more
Hugging Face is exploring a sale that would value the AI model hub at $13 billion or more, Business Insider reported on August 23. No buyer is named and no deal is agreed. The site hosts more than 2 million models.
- Drew Breunig — expensive models ended the free lunch in AI coding
Drew Breunig argues that Claude Fable 5 broke the habit of waiting for a cheaper model to solve your problems. He now sends design work to Fable and rote coding to GLM 5.2, which he puts at about one-ninth of Fable's cost.
- GLM-5.3 and Kimi K3 root an Amazon Fire tablet — $266 of AI, one 2022 CVE
Kimi K3, GLM-5.2 and GLM-5.3 built a working root exploit for an Amazon Fire HD 10 in a write-up that cost $266.15 in AI billing. Claude and ChatGPT refused the exploit work; the Chinese models finished it.
- Wes Roth — 'Ilya Sutskever new Superintelligence model will change EVERYTHING'
Wes Roth's August 23 video takes on the first model expected from Safe Superintelligence, Ilya Sutskever's lab. SSI itself has announced nothing: ssi.inc is still a mission statement and a hiring page.
- MoneyPrinterTurbo v1.3.5 — Claude joins the one-click short-video maker
MoneyPrinterTurbo v1.3.5 adds Anthropic Claude as a script model, MiniMax and Fish Audio voices, and two text-to-video sources. The MIT-licensed tool turns one keyword into a finished vertical short and has 115,138 GitHub stars.
- LiteLLM v1.98.0 — reserved capacity gets flat-cost billing, not per-token
LiteLLM v1.98.0 adds provisioned-throughput billing, so a deployment on reserved capacity carries a flat cost instead of a per-token charge. The release also ships shadow evals for the auto-router and a per-key prompt caching switch.
- Ray 2.58.0 — KV-cache-aware routing lands for LLM serving
Ray 2.58.0 finishes the KV-cache and token-aware request routing that Ray Serve LLM previewed in 2.57. The router tokenizes inside its own ingress replica and passes tokens out-of-band, so the engine never tokenizes twice.
- Grok 4.6 on Google's agent platform — xAI's flagship arrives in Model Garden
Grok 4.6 is now available on the Google Enterprise Agent Platform through Model Garden. xAI's flagship keeps its 500K-token context window and four reasoning levels, at $2 per million input tokens and $6 per million output.
- OpenViking v0.4.16 — agents can run Skills hosted on another server
OpenViking v0.4.16 lets VikingBot find, cache and run Skills stored on a remote OpenViking server, adds a per-user memory extraction policy for admins, and removes the experimental Resource Relations API.
- Claude Code 2.1.239 — a proxy bug that doubled Bedrock API calls is fixed
Claude Code 2.1.239 fixes a bug where Bedrock streaming behind proxies that strip the Content-Type header silently doubled billed API calls. The release also adds /claude-api upgrade, which migrates Python projects to anthropic 1.x.
- Lucian Ghinda — a week of reaching for Codex instead of Claude Code
Lucian Ghinda spent a week using OpenAI's Codex more than Claude Code. He reports that Codex writes simpler Ruby with fewer comments, while Claude Code anticipates edge cases and fits his CLI-based Jira workflow better.
- SGLang v0.5.18 — cold starts get 2.38x faster, seven model families land
SGLang v0.5.18 stages model weights from storage while CUDA graphs capture, so a Qwen3-32B server on an H100 starts in 35.6 seconds instead of 84.8. Seven new model families get serving support. 710 pull requests from 212 contributors.