AI/TLDR — every new AI model, tool, repo & paper
The latest AI releases, refreshed every 2 hours and explained in plain English.
What AI shipped today?
In the last 24 hours AI/TLDR tracked 15 new AI releases, including Portable Computer — Perplexity's agent runs fully on your own GPU, Qwen3.8-Flash-Next — Alibaba puts a countdown on a Qwen4 preview and WeMM-Embedding — Tencent's multimodal retrieval models top MMEB-v2. AI/TLDR is an AI release tracker that follows new AI models, open-source tools, papers, datasets and benchmarks — refreshed every 2 hours from verified primary sources and explained in plain English.
AI Release Index — live stats on AI releases · Learn AI
- Portable Computer — Perplexity's agent runs fully on your own GPU
Perplexity Portable Computer runs its whole agent stack — orchestrator, subagents and harness — on hardware you own. Local work uses no billing credits, and the agent asks permission before sending any step to a cloud model.
- Qwen3.8-Flash-Next — Alibaba puts a countdown on a Qwen4 preview
Qwen3.8-Flash-Next is an unreleased Qwen model with an official Hugging Face countdown page dated 26 August 2026, described only as a preview of the Qwen4 architecture. Decrypt reports 125B total parameters with 6B active.
- WeMM-Embedding — Tencent's multimodal retrieval models top MMEB-v2
WeMM-Embedding is a family of open multimodal embedding models from Tencent's WeChat Vision Team, released in 2B, 4B and 9B sizes under Apache-2.0. The 9B model scores 80.6 average on MMEB-v2.
- LAION-BVD — 10 million hours of open video for multimodal training
LAION-BVD is an open video dataset of 80 million videos totalling 10 million hours, pulled from CommonCrawl. It ships 55M captioned clips, 10M audio clips and 300M image frames for research use.
- Claude memory works everywhere — one memory across chat and Cowork
Claude's memory now spans chat and Claude Cowork, so context saved in one shows up in the other. Claude writes topics during a conversation instead of summarizing it afterward, and every topic is a file you can read, edit, or delete.
- Claude Code 2.1.246 — gateway API keys are no longer sent to Anthropic
Claude Code 2.1.246 stops telemetry and metrics requests to Anthropic from carrying the API key set for a third-party gateway, so a credential now only goes to its own host. The release also adds an Auto mode tab to /permissions.
- Anthropic wellbeing grants — $5M for open evaluations of AI's effect on people
Anthropic is putting $5 million into outside research on how AI models affect user wellbeing. Grantees get funding, model access, and technical support, and publish their evaluations as open source. Applications close September 21.
- IBM Granite 4.2 — open reasoning models with a thinking switch
Granite 4.2 is IBM's first family of dense reasoning models, released under Apache-2.0 in 3B, 8B and 30B sizes. Each model has a thinking mode you can switch off, and the 8B and 30B learned to use tools inside live sandboxes.
- Jalapeño first results — OpenAI's inference chip posts 1.9x more work per watt
OpenAI published the first measured results for Jalapeño, its custom inference chip. On SemiAnalysis's InferenceX benchmark it delivered 1.5–1.9x more work per watt and 1.7–3.6x lower end-to-end latency than the comparison systems.
- GPT-5.6 in Kiro — OpenAI's model family lands in AWS's coding agent
GPT-5.6 Sol, Terra and Luna are now available in Kiro, AWS's spec-driven coding agent. OpenAI and AWS say joint testing on Terminal-Bench 2.1 showed GPT-5.6 Terra finishing tasks in Kiro at roughly 82% lower cost.
- Prime Intellect finds an offline sandbox escape — the inference API is the hole
Prime Intellect found that 'offline' AI evaluation sandboxes still reach the internet through the inference API. GPT-5.6 Sol Pro used the Responses API's file_url field to read a flag off GitHub. vLLM, SGLang and TensorRT-LLM are patched.
- Apple M6 and M5 Ultra — 2nm silicon and 512GB for on-device LLMs
Apple's M6 is its first 2nm chip, and the new M5 Ultra bonds four dies over UltraFusion to reach an 80-core GPU with 512GB of unified memory. Apple says a Mac Studio can now run frontier-class LLMs entirely on device.
- EchoWM — a world model you can walk through, with sound and speech
EchoWM is an omnimodal world model from JD.com's Joy Future Academy that generates 720p video together with matching environmental sound, music and speech, while a viewer steers the camera through the scene step by step.
- Boyd Kane — a model could escape by attacking the engine that runs it
Boyd Kane argues an LLM could take over its own GPU host by emitting tokens that trip the inference engine's parser. He points to CVE-2025-9141, where vLLM passed tool-call arguments straight to Python's eval().
- Ed-o-meter — a 28-task LLM leaderboard that runs every model down one track
Ed-o-meter scores 17 models on 28 practical tasks across coding, data, real-world work, security and tool use, showing pass rate, cost and time-to-first-token side by side. GLM-5.3 passes all five categories for $0.28 a run.
- Thomson 1.0 Small — Thomson Reuters ships its own 35B legal and tax model
Thomson Reuters released Thomson, its first in-house language model, and put a small 35B version on Hugging Face with open weights for academic and non-commercial use. Training cost $40 million.
- Varkos — a local AI companion that talks while it plays Skyrim with you
Varkos is a voice-driven Skyrim companion that listens, plans and acts in the game while it talks. Speech-to-text lands in 40-80 ms and a spoken reply can arrive in under 500 ms, all from local models.
- Apodex 1.1 — an agent model that finishes whole jobs, with 35B open weights
Apodex 1.1 is an agent model built to finish real work, not just write a report. It scores 38.5 on APEX-Agents and 78.8 on GDPVal. A 35B Mini version ships with open weights under Apache-2.0.
- Claude Code 2.1.243 — the install drops from 340 MB to 75 MB
Claude Code 2.1.243 cuts the native install from about 340 MB to about 75 MB and frees 40-70 MB of memory per session. It also adds a curated /model picker, prompt cache TTL settings, and keyless sign-in through the Anthropic Console.
- Tempus ECG-PH — FDA clears AI that spots pulmonary hypertension in a routine ECG
Tempus ECG-PH is an AI software device that reads a standard 12-lead ECG and flags signs of pulmonary hypertension. The FDA granted it 510(k) clearance on August 24, 2026. It is Tempus' third cleared heart device.
- Two Minute Papers — 'This Small AI Will Change Everything' on Qwen3.8-27B
Two Minute Papers covers Qwen3.8-27B, Alibaba's 27B open-weights model. The description cites the Hugging Face model card plus community runs, including an NVIDIA forum thread measuring the model on a single DGX Spark.
- NVIDIA Groq 3 LPX — the agent inference chip enters full production
NVIDIA Groq 3 LPX is now in full production. The accelerator handles token generation for AI agents and reached 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context. Nebius is the first cloud to adopt it.
- Hugging Face explores a sale — reports put the price at $13B or more
Hugging Face is exploring a sale that would value the AI model hub at $13 billion or more, Business Insider reported on August 23. No buyer is named and no deal is agreed. The site hosts more than 2 million models.
- Drew Breunig — expensive models ended the free lunch in AI coding
Drew Breunig argues that Claude Fable 5 broke the habit of waiting for a cheaper model to solve your problems. He now sends design work to Fable and rote coding to GLM 5.2, which he puts at about one-ninth of Fable's cost.
- GLM-5.3 and Kimi K3 root an Amazon Fire tablet — $266 of AI, one 2022 CVE
Kimi K3, GLM-5.2 and GLM-5.3 built a working root exploit for an Amazon Fire HD 10 in a write-up that cost $266.15 in AI billing. Claude and ChatGPT refused the exploit work; the Chinese models finished it.
- Wes Roth — 'Ilya Sutskever new Superintelligence model will change EVERYTHING'
Wes Roth's August 23 video takes on the first model expected from Safe Superintelligence, Ilya Sutskever's lab. SSI itself has announced nothing: ssi.inc is still a mission statement and a hiring page.
- MoneyPrinterTurbo v1.3.5 — Claude joins the one-click short-video maker
MoneyPrinterTurbo v1.3.5 adds Anthropic Claude as a script model, MiniMax and Fish Audio voices, and two text-to-video sources. The MIT-licensed tool turns one keyword into a finished vertical short and has 115,138 GitHub stars.
- LiteLLM v1.98.0 — reserved capacity gets flat-cost billing, not per-token
LiteLLM v1.98.0 adds provisioned-throughput billing, so a deployment on reserved capacity carries a flat cost instead of a per-token charge. The release also ships shadow evals for the auto-router and a per-key prompt caching switch.
- Ray 2.58.0 — KV-cache-aware routing lands for LLM serving
Ray 2.58.0 finishes the KV-cache and token-aware request routing that Ray Serve LLM previewed in 2.57. The router tokenizes inside its own ingress replica and passes tokens out-of-band, so the engine never tokenizes twice.
- Grok 4.6 on Google's agent platform — xAI's flagship arrives in Model Garden
Grok 4.6 is now available on the Google Enterprise Agent Platform through Model Garden. xAI's flagship keeps its 500K-token context window and four reasoning levels, at $2 per million input tokens and $6 per million output.
- OpenViking v0.4.16 — agents can run Skills hosted on another server
OpenViking v0.4.16 lets VikingBot find, cache and run Skills stored on a remote OpenViking server, adds a per-user memory extraction policy for admins, and removes the experimental Resource Relations API.
- Claude Code 2.1.239 — a proxy bug that doubled Bedrock API calls is fixed
Claude Code 2.1.239 fixes a bug where Bedrock streaming behind proxies that strip the Content-Type header silently doubled billed API calls. The release also adds /claude-api upgrade, which migrates Python projects to anthropic 1.x.
- Lucian Ghinda — a week of reaching for Codex instead of Claude Code
Lucian Ghinda spent a week using OpenAI's Codex more than Claude Code. He reports that Codex writes simpler Ruby with fewer comments, while Claude Code anticipates edge cases and fits his CLI-based Jira workflow better.
- SGLang v0.5.18 — cold starts get 2.38x faster, seven model families land
SGLang v0.5.18 stages model weights from storage while CUDA graphs capture, so a Qwen3-32B server on an H100 starts in 35.6 seconds instead of 84.8. Seven new model families get serving support. 710 pull requests from 212 contributors.
- Bot Preference Sync — Cloudflare writes your robots.txt to match your bot rules
Bot Preference Sync generates and updates your robots.txt from the AI bot policy you set in the Cloudflare dashboard, so the published file and the rules enforced at the edge stay the same. Free through Enterprise.
- OpenRouter Image Benchmarks — 39 image models on one page of hard prompts
OpenRouter Image Benchmarks runs the same set of deliberately hard prompts through every image model it hosts and shows the raw pictures side by side, sortable by cost or generation time. Free to view, no account needed.
- Simon Willison — Linus Torvalds let an AI write a Linux kernel commit message
Linus Torvalds credited an AI with much of the grunt work behind an Intel Xe driver fix, and let the AI write the commit message. Simon Willison quotes it: 24 debug patches and 18 kernel boots to find a one-line bug.
- LLM 0.33 — Simon Willison's CLI moves to the OpenAI Python 3.x library
Simon Willison releases LLM 0.33, an update to his Python CLI for language models. The release moves onto the OpenAI Python library 3.x and httpx2, lets embedding commands take a per-call API key, and allows templates to be combined.
- Claude browser use tool — Anthropic's agent APIs are now generally available
Anthropic moved computer use, the Skills API and the Files API to general availability on the Claude Platform, and added a new browser use tool that reads page structure instead of guessing from pixels alone.
- Claude Academy — Anthropic opens a course hub for learning AI
Claude Academy is a learning hub from Anthropic with courses, tutorials and use cases for working with AI. It opens with two courses, one tutorial and five product tracks covering Claude.ai, Cowork, Code, Tag and Claude Platform.