New AI Model Releases — LLMs & Open Weights
Every new AI model worth knowing — frontier and open-weight releases, explained in plain English with the context window, parameters and benchmarks that matter.
188 releases tracked
- Qwen3.8-Max — Alibaba's 2.4T flagship ships official with a full benchmark table
Qwen3.8-Max lands as Alibaba's official 2.4T-parameter, 95B-active MoE flagship with published benchmarks and $2/$6 per-million-token pricing.
- Seedance 2.5 — ByteDance's video model doubles to 30 seconds per generation
ByteDance's Seed lab doubles single-take video generation to 30 seconds with 50-asset multimodal referencing.
- DeepHealth Breast Ultrasound — FDA clears RadNet AI for automated lesion reads
AI that reads breast ultrasounds now has FDA sign-off, and RadNet is rolling it out across its US imaging centers.
- Inkling-Small — Thinking Machines' 276B open model matches Inkling at 1/4 the size
A 276B MoE with 12B active parameters that ships full Apache-2.0 weights and matches its 4x-larger sibling.
- Gemini Robotics ER 2 — the planning brain that watches video and coordinates robots
DeepMind's embodied-reasoning brain plans multi-step robot work, watches the video, and orchestrates multiple robots at once.
- MiniMax H3 — open-weights video model does 2K, 15s, and native stereo sound
A single open-weights video model does 2K clips, stereo audio, editing, and motion transfer up to 15 seconds long.
- DeepSeek V4-Flash goes official — 0731 hits 82.7 on Terminal Bench 2.1
DeepSeek graduates V4-Flash from preview to official with a re-post-trained checkpoint tuned for agents.
- GPT-5.6 Luna cut 80%, Terra 20% — OpenAI drops API prices three weeks after launch
OpenAI drops GPT-5.6 Luna 80% and Terra 20% three weeks after launch, passing on 20% serving-cost gains.
- Gemini Robotics 2 — DeepMind's new humanoid model controls the full body
DeepMind's new humanoid brain walks, crouches, and ties knots — and coordinates two robots at once.
- Grok Voice Think Fast 2.0 — xAI's new voice model ships at 0.70s time-to-first-audio
xAI's next-gen voice model tops the Artificial Analysis Speech-to-Speech leaderboard at sub-second latency.
- Lyria 3.5 — Google DeepMind's new music model in free Flow Music
Google DeepMind's new music model — richer melodies, clearer vocals, longer songs in the free Flow Music app.
- Microsoft Mage-VL — 4B codec-native multimodal cuts video tokens by 75%
Mage-VL treats a video like a codec — keep anchor frames, drop most predicted-frame patches, and cut visual tokens by 75%.
- KAT-Coder-V2.5-Dev — Kwaipilot's 35B open-weight agentic coding MoE
Kwaipilot opens the weights of a 35B/3B-active MoE tuned to act inside real code repositories, not just autocomplete a snippet.
- Kimi K3 open weights — Moonshot drops the 2.8T MoE on Hugging Face
The largest open-weight AI model in history — 2.8T Kimi K3 is now free to self-host.
- Microsoft MAI-Cyber-1-Flash — 96% on CyberGym at half the price of the GPT-5.4 stack
Microsoft's new security-specialist model plugs into MDASH and finds bugs cheaper and better than a GPT-5.4 stack.
- Inflect-Micro-v2 — 9.36M-parameter local text-to-speech under 40 MB
A complete English text-to-speech model that fits in 37.5 MB and beats larger systems 66% of the time in blind human tests.
- Apertus 1.5 — Switzerland's fully open 8B/70B goes multimodal
Switzerland's fully open Apertus grows eyes and ears, keeps its Apache-2.0 promise.
- FLUX 3 x mimic — the FLUX 3 backbone drives factory robots at 101 ms
A FLUX 3 spin-off that lets one on-robot GPU pilot a real factory arm in near real time.
- Claude Opus 5 — Anthropic's new Opus nears Fable 5 at half the price
Anthropic's new Opus tier comes close to Fable 5 while holding the Opus 4.8 price.
- FLUX 3 — Black Forest Labs' multimodal video, image, and robotics model
One Black Forest Labs model that generates video with audio, edits images, and drives real robots at Audi.
- Nanbeige4.2-3B — 3B looped-transformer open-weight agent hits 63.6 on SWE-Bench Verified
A 3B open-weight agent whose Looped Transformer reuses layers to match models 3–4× its size.
- Cisco Antares — 350M and 1B open-weight models that hunt code vulnerabilities
Small on-prem models that read a CWE, walk your repo, and tell you which files the vulnerability is hiding in.
- Microsoft Mage-Flow — 4B native-resolution image model that keeps up with 20B systems
A 4B image model from Microsoft Research that generates or edits at any resolution and matches open systems five times its size.
- Solar Open 2 — Upstage's 250B/15B open-weight MoE built for agentic use
Korea's Upstage ships a 250B/15B open-weight MoE with 1M context and a hybrid-attention stack, aimed at agentic work.
- Poolside Laguna S 2.1 — 118B open-weight coding MoE with 8B active
Poolside ships Laguna S 2.1 — a 118B/8B-active open-weight coding MoE with a 1M-token context, pitched as the West's answer to DeepSeek and Qwen.
- Gemini 3.6 Flash — Google's workhorse plus Flash-Lite and Flash Cyber
Google ships three Flash-tier Gemini variants tuned for coding, high-throughput agents, and cybersecurity.
- Qwen-Image-3.0 — Alibaba's third-gen image model ships without weights
Alibaba's third-gen image model takes 4,500-token prompts and renders dense text — but ships without weights, license, or benchmarks.
- Qwen-Audio-3.0-TTS — Alibaba's TTS hits #1 on Artificial Analysis
Alibaba's new hosted TTS ships in Flash and Plus tiers, spans 16 languages, and takes #1 on the Artificial Analysis speech leaderboard.
- Qwen3.8-Max Preview — Alibaba's 2.4T multimodal flagship arrives, weights promised
Alibaba announces Qwen3.8-Max, a 2.4T multimodal preview, and promises open weights soon.
- NVIDIA Cosmos 3 Edge — 4B world model that runs physical AI on Jetson
NVIDIA's Cosmos family gets a small, on-device sibling built for real-time robotics.
- OvisOCR2 — 0.8B Alibaba model tops OmniDocBench and beats pipeline OCR
Alibaba's 0.8B end-to-end document parser sets state of the art on OmniDocBench v1.6.
- Nemotron 3 Embed — NVIDIA's open 8B embedder takes #1 on RTEB
NVIDIA ships an open 8B embedder that takes #1 on RTEB, plus two efficient 1B variants for production RAG.
- Kimi K3 — Moonshot's 2.8T flagship with 1M context lands on web, app, and API
Moonshot AI ships Kimi K3 today, jumping to a 2.8T mixture-of-experts with a 1M-token context and native vision.
- Descartes — Hemispheric's frontier NeuroAI model for decoding EEG brain signals
A 6B-parameter AI model that converts 15-minute EEG recordings into quantitative brain-health diagnostics.
- NeuroVFM — brain-scan AI outperforms GPT-5 on clinical triage
Michigan Medicine's 5M-scan neuroimaging AI trained on uncurated hospital data outperforms GPT-5 on triage at 24× lower cost.
- Boogu-Image-0.1 — 10B open image model trained on 208M images for ~$400K
Apache-2.0 10B image generation model reports near closed-source quality after training on 208M images for about $400K in compute.
- MonkeyOCRv2 — visual-text foundation model for document AI
MonkeyOCRv2 pretrains a visual-text foundation model on 113M multilingual document images.
- Inkling — Thinking Machines' first open-weights 975B/41B multimodal MoE
Mira Murati's Thinking Machines Lab ships its first foundation model, and puts the full weights on HuggingFace under Apache 2.0.
- Xiaomi-Robotics-U0 — open 38B unified embodied world model
38B open autoregressive foundation model that generates images, embodied scenes and robot video from one shared tokenizer.
- Bonsai 27B — first 27B-class LLM to run on a phone at ~4 GB
Bonsai 27B is the first 27B-class open-weight model small enough to run on a phone — 3.9 GB of weights, 262K context, Apache-2.0.
- Seedream 5.0 Pro — ByteDance's image model reasons before it draws
ByteDance's Seedream 5.0 Pro thinks about the prompt before it draws, then hits 2K with clean text in 14 languages.
- GPT-5.6 goes public — Sol, Terra, and Luna clear the White House gate
OpenAI's frontier model line goes broad after the government finishes its second look.
- Muse Spark 1.1 — Meta MSL opens its first paid API at $1.25 / $4.25 per million tokens
Meta finally opens Muse Spark to outside developers — with a price tag under Claude and GPT.
- Grok 4.5 — xAI's Opus-class flagship at $2/$6 per million tokens
xAI's new coding flagship ties GPT-5.5 on Terminal-Bench at a quarter of the price of Fable 5 and Opus.
- OpenAI GPT-Live — full-duplex voice model now powers ChatGPT Voice
OpenAI GPT-Live is a full-duplex voice model that lets ChatGPT listen and speak at the same time.
- SWE-1.7 — Cognition's coding model runs on Devin at 1000 tok/s via Cerebras
SWE-1.7 is Cognition's new coding model — near GPT-5.5 on agentic coding, running at 1000 tokens per second on Devin.
- Robostral Navigate — Mistral's first embodied model, 8B, single RGB camera
Mistral's first embodied model steers wheeled, legged, and flying robots from a single RGB camera and a plain-language instruction.
- Claude Fable 5 access extended again — Anthropic pushes subscription window to July 19
Anthropic gives Fable 5 subscribers another week before usage credits kick in.
- Cohere Transcribe Arabic — 2B open-weight ASR beats Whisper on dialect and code-switching
The open-source Arabic speech model that finally beats Whisper on dialect audio.
- Meta Muse Image — MSL's first image model plans, calls tools, and self-refines
Meta Superintelligence Labs' first image model plans, calls tools, and self-refines like a reasoning LLM.
- Ternlight — 7 MB embedding model that runs in the browser
A 7 MB sentence-embedding model that runs in the browser at ~5 ms per query, with no API call.
- xAI Grok Voice — 21 new flagship voices with speech tags for pacing
Grok Voice now ships 21 new multilingual flagship voices plus inline speech tags for pacing and whispering.
- OpenAI gpt-realtime-2.1 — voice model gains reasoning-effort dial and a mini variant
gpt-realtime-2.1 adds a reasoning-effort dial and better silence/interruption handling to OpenAI's Realtime API.
- Tencent Hunyuan Hy3 — 295B/21B MoE open-sourced with 256K context
Tencent open-sources its Hunyuan flagship: 295B total, 21B active, 256K context, Apache 2.0.
- Leanstral 1.5 — Mistral's updated Lean 4 formal-proof model
Mistral's Lean 4 theorem-prover saturates miniF2F, tops PutnamBench, and ships as an Apache-2.0 119B/6.5B MoE.
- TabFM — Google's zero-shot foundation model for tabular data
A pretrained-once foundation model that skips XGBoost's tuning ritual for tabular classification and regression.
- Gemini Omni Flash + Nano Banana 2 Lite — Google's new video and image models
Google ships two preview models: a 10-second video generator priced per second, and a 4-second image generator priced per shot.
- Claude Sonnet 5 — Anthropic's new agentic Sonnet at Opus-class quality
Anthropic's new mid-tier Sonnet 5 lands with a 1M-token context, adaptive thinking, and $3/$15 pricing — Opus-class quality without Opus pricing.
- Agents-A1 — Shanghai AI Lab 35B MoE matches trillion-parameter agents
Open-weight 35B agent from Shanghai AI Lab posts SOTA on SEAL-0, IFBench, and FrontierScience-Research.
- LongCat-2.0 — Meituan's 1.6T open-source MoE for agentic coding
Meituan's first frontier-tier open-source coding model — 1.6T parameters with 48B active, trained without a single NVIDIA chip.
- GPT-5.6 — OpenAI previews Sol, Terra, and Luna tiers
OpenAI's new generation splits into three named tiers and adds an ultra mode that wires subagents into the flagship model.
- Ornith 1.0 — open-weight coding models that learn their own RL scaffold
Open-weight agentic coding models that learn to write their own RL scaffold instead of relying on a fixed harness.
- Gemini 3.5 Flash gets Computer Use — native browser, mobile, and desktop agents
Gemini 3.5 Flash can now see a screen and click, type, and scroll on its own through a single built-in tool.
- Krea 2 — open-weight 12B image model with 2-second Turbo variant
Krea AI open-sources a 12B Diffusion Transformer image model with a Turbo variant that draws 2K in two seconds.
- Mistral OCR 4 — 170-language document model with bounding boxes and confidence scores
Mistral's new document model returns structured pages with boxes, block types, and per-word confidence at $4 per 1,000.
- Baidu Unlimited-OCR — 3B vision model parses long documents in one pass
Baidu's open 3B OCR model swaps standard attention for R-SWA so it can transcribe dozens of pages without the usual KV-cache blowup.
- PP-OCRv6 — PaddlePaddle ships 50-language OCR family from 1.5M to 34.5M params
PaddlePaddle's PP-OCRv6 is a three-tier OCR family — Tiny 1.5M to Medium 34.5M — that recognises 50 languages and beats PP-OCRv5_server.
- Sakana Fugu — multi-agent orchestration model that matches Fable 5 on quality
Sakana AI ships Fugu, a single API that routes each request to a pool of frontier models and verifies the answer before returning it.
- MolmoMotion — Ai2's language-guided 3D motion forecasting models
MolmoMotion predicts how points on objects move in 3D from a video frame and a text instruction, with weights, a 1.16M-video dataset, and a benchmark.
- Qwen-Robot Suite — Alibaba's three foundation models for robots
Three open foundation models from Alibaba's Qwen team that move robots, drive them around, and predict what happens next.
- VibeThinker-3B — Weibo's 3B reasoning model hits 80.2% on LiveCodeBench v6
Sina Weibo's 3B model finetuned from Qwen2.5-Coder-3B, MIT-licensed, scoring 94.3 on AIME26 and 80.2 on LiveCodeBench v6.
- Grok Imagine Video 1.5 — xAI's image-to-video model goes GA at $0.14/sec 720p
xAI's image-to-video model — the engine behind Grok Imagine's video clips — is now generally available as a pay-per-second API.
- GLM-5.2 — Z.ai's new flagship coding model with 1M context
Z.ai's new coding flagship lands first inside the GLM Coding Plan, with API, chatbot, and open weights set for next week.
- Kimi K2.7-Code — Moonshot's 1T MoE Coding Model Beats Claude Opus on MCPMark
Moonshot's coding-specialized fork of Kimi K2.6, with faster reasoning and a higher MCP tool-use score than Claude Opus 4.8.
- Decart Ships Oasis 3 — First API-Accessible Interactive World Model for Physical AI Streams Three Synchronized 768×512 Camera Views at 22 FPS With <200ms End-to-End Latency on NVIDIA HGX B200, Priced at $0.02 per Second of Simulation
Decart's Oasis 3 lets robotics and AV teams stream three synchronized photorealistic camera views in real time from a text prompt, action-conditioned, via API.
- Amap Open-Sources ABot-Earth 0.5 — Alibaba's 3D Native World Model Generates Kilometer-Scale 3D Gaussian-Splatting City Scenes From a Single Satellite Image or Text Prompt in About 10 Minutes per km² on a Consumer GPU
Alibaba's Amap turns one satellite image into an interactive 3D city in roughly 10 minutes per square kilometer.
- Google Ships DiffusionGemma — Apache-2.0 26B/3.8B-Active Mixture-of-Experts That Denoises 256 Tokens in Parallel via Discrete Block Diffusion, Hits 1,000+ Tokens/Sec on H100 and 700+ on RTX 5090 While Posting 77.6% MMLU Pro, 73.2% GPQA Diamond, and 69.1% LiveCodeBench v6
Google opens a 26B-parameter Gemma that denoises 256 tokens at once instead of generating them one by one.
- Kuaishou Open-Sources Keye-VL-2.0-30B-A3B — Apache-2.0 Mixture-of-Experts Vision-Language Model Activates 3B Parameters Per Token, Lands 256K Context With DeepSeek Sparse Attention for Lossless Long-Video Reasoning, Beats Qwen3-VL-235B on LongVideoBench at 74.1, and Tops 70.1 mIoU on QVHighlights-TimeLens
A 30B/3B-active MoE that pushes long-video understanding past dense 235B baselines, all under Apache-2.0.
- Cohere Ships North Mini Code 1.0 — Apache-2.0 Sparse 30B/3B-Active Mixture-of-Experts Coding Model With a 256K Context, 64K Max Output, 83.2% Pass@1 on SWE-Bench Verified, 63% on Terminal-Bench v2, and 2.8× the Output Throughput of Devstral Small 2
Cohere's first model aimed squarely at developers — a 30B/3B-active MoE coding model under Apache 2.0.
- Anthropic Ships Claude Fable 5 and Claude Mythos 5 — New Flagship Tops Hebbia's Finance Benchmark and Cognition's FrontierCode Eval at $10/$50 per Million Tokens, With Mythos 5 Restricted to Project Glasswing Partners and Select Biology Researchers
Fable 5 ships to everyone today; Mythos 5 stays gated to Glasswing partners and biology researchers.
- Google DeepMind Ships Gemini 3.5 Live Translate — Audio Model Streams Near Real-Time Speech-to-Speech in 70+ Languages With Preserved Intonation, SynthID Watermarks, and Public Preview on Gemini Live API, AI Studio, Meet, and Google Translate
First production speech-to-speech model that translates continuously instead of taking turns.
- Xiaomi Ships MiMo-V2.5-Pro-UltraSpeed — FP4-Quantized 1.02T MoE Hits ~1,200 Tokens/Sec on Stock 8-GPU Nodes via Block-Diffusion 'DFlash' Speculative Decoding, MIT-Licensed FP4-DFlash Checkpoint Lands on Hugging Face
An FP4 backbone plus a block-diffusion drafter pushes Xiaomi's 1T MiMo MoE past 1,000 tokens per second on a stock 8-GPU node.
- NVIDIA Nemotron 3.5 ASR Streaming 0.6B — Open-Weights Multilingual Speech Model Sustains ~17× More Concurrent Streams on H100, Covers 40 Locales From One Checkpoint
A 600M-param open ASR model that streams 40 locales from one checkpoint and squeezes ~17× more concurrent voice sessions onto an H100.
- Google Ships Gemma 4 QAT Checkpoints — New Mobile Quantization Format Brings Gemma 4 E2B Under 1GB of RAM, Q4_0 Weights Land for E2B, E4B, 12B, and the 26B MoE on Hugging Face
Gemma 4 shrinks to under 1GB on phones via a custom 2-bit mobile quantization format trained with QAT.
- OpenAI Updates GPT-Rosalind With GPT-5.5 Reasoning — New LifeSciBench Tops Internal Evals, Two Codex Life-Sciences Plugins Land, and Novo Nordisk Joins the Trusted-Access Roster
OpenAI's life-sciences specialist gets a GPT-5.5 brain, two Codex plugins for bio workflows, and Novo Nordisk on the partner list.
- Magenta RealTime 2 — Google's Open-Weights Live Music Model Lands With ~200ms Control Latency, Native Apple Silicon C++/MLX Inference, and Both 2.4B and 230M Variants on CC BY 4.0
An open-weights live music model that responds to MIDI, text, and audio in ~200ms and runs entirely on your MacBook.
- Ideogram 4.0 — 9.3B Open-Weight Text-to-Image Foundation Model Lands With Apache-2.0 Inference Code, Native 2K Resolution, and Structured JSON Prompting That Tops DesignArena's Open-Weight Leaderboard
Ideogram's first open-weight image model — 9.3B params, native 2K output, JSON-controlled layouts, and SOTA text rendering for design.
- Gemma 4 12B — Google's First Unified, Encoder-Free Multimodal Decoder Streams Audio, Video, and Image Patches Straight Into an 11.95B Apache-2.0 LLM With 256K Context
A 12B open multimodal model that lets raw audio and video patches skip the encoder and hit the LLM directly.
- H Company Open-Sources Holo3.1 — Computer-Use VLM Family at 0.8B, 4B, 9B, and 35B-A3B Hits 74.2% OSWorld, Jumps AndroidWorld From 67% to 79.3%, and Ships FP8, NVFP4, and Q4 GGUF Checkpoints Under Apache 2.0
An Apache 2.0 computer-use VLM family that runs locally on consumer GPUs and now drives mobile as well as desktop.
- NVIDIA Nemotron 3 Ultra — 550B Mamba-Transformer Mixture-of-Experts With 55B Active Parameters Tops US Open-Weights Leaderboard at 48 on Artificial Analysis Intelligence Index, Ships June 4 on Hugging Face, OpenRouter, ModelScope, and build.nvidia.com
NVIDIA's biggest open-weights model: a 550B Mamba-Transformer MoE that beats every other US open model on intelligence and inference speed.
- WindBorne WeatherMesh-6 — AI Weather Model Cuts Ensemble-Mean RMSE Up to 38% vs ECMWF, Produces Hourly 3 km Forecasts From a 128-Member Ensemble in Latent Space
An AI weather model that beats ECMWF's flagship physics forecast by up to 38% on ensemble RMSE — and refreshes hourly.
- Microsoft Aion 1.0 — On-Device Windows AI Lineup Debuts at Build 2026 With Aion 1.0 Instruct in Edge Insider Today and a 14B Aion 1.0 Plan Reasoning + Tool-Calling Model Shipping In-Box
Microsoft ships a Windows-native AI model line: a small Instruct SLM for everyday text tasks and a 14B agentic Plan model that runs in-box.
- MiniMax M3 — Open-Weight Frontier Coding, 1M-Token Context, and Native Multimodality on a New Sparse-Attention Architecture
M3 packages frontier coding, a million-token context, and native multimodality into one open-weight model.
- Microsoft Launches Seven New MAI Models at Build 2026 — MAI-Thinking-1 Reasoning, MAI-Code-1-Flash 5B for Copilot, MAI-Image-2.5 With a Flash Variant, MAI-Voice-2 Across 15 Languages, MAI-Transcribe-1.5 Across 43
Microsoft AI's biggest first-party launch yet: a reasoner, a coding flash model, image + voice updates, and a speech-to-text engine claimed to be 5x faster than competing models.
- NVIDIA Cosmos 3 — Open Mixture-of-Transformers Omni-Model for Physical AI Lands at Computex With Nano (16B) and Super (65B) Variants, OpenMDW 1.1 Weights, and a New Coalition With Runway, Black Forest Labs, and Skild AI
An open frontier omni-model that can reason about, simulate, and act in the physical world.
- JetBrains Mellum2 — 12B Mixture-of-Experts With 2.5B Active Parameters Opens Up Under Apache 2.0, Ships Base, Instruct, and Thinking Variants Trained on ~10.6T Tokens
A fast, open 12B MoE designed for routing, RAG, and sub-agents inside coding workflows.
- Grok Build 0.1 Lands on the xAI API in Public Beta — 256K Context, $1/$2 per 1M Tokens, 100+ Tokens/Sec for Agentic Coding
xAI's coding-focused model goes from CLI-only to a public-beta API at $1 in, $2 out per million tokens.
- Qwen Ships Qwen-VLA — Unified Vision-Language-Action Generalist Built on Qwen3.5-4B With a 1.15B DiT Flow-Matching Action Decoder, 97.9% on LIBERO and 76.9% Real-World ALOHA
Qwen's vision-language stack now talks to robots end-to-end.
- Liquid AI Ships LFM2.5-8B-A1B — On-Device MoE With 1B Active Parameters, 38T Training Tokens, 128K Context, and Explicit Chain-of-Thought
8B-total / 1B-active mixture-of-experts trained on 38 trillion tokens, tuned for on-device agentic work with explicit chain-of-thought.
- Stepfun Releases Step 3.7 Flash — 198B Vision-Language MoE With 11B Active Parameters, 56.3% SWE-Bench Pro, 256K Context, and Three Reasoning Levels
198B-parameter vision-language MoE tuned for high-frequency agentic workloads — Apache 2.0, 256K context, and three reasoning levels you can dial up or down.
- NVIDIA LocateAnything-3B — Parallel Box Decoding Vision-Language Grounder Hits 12.7 Boxes/Sec on H100, ~10× Faster Than Qwen3-VL, Trained on 138M Queries Across 785M Boxes
LocateAnything decodes full bounding boxes in one shot, getting NVIDIA a vision-language grounder that is roughly 10× faster than Qwen3-VL.
- Claude Opus 4.8
Anthropic's flagship gets sharper agentic judgment, honest error-flagging, and a parallel-subagent dynamic-workflow tool.
- LLaVA-OneVision-2 — Open 8B Vision-Language Model Reads Video as a Codec Bit-Cost Stream, Hits 74.9 JumpScore mAP vs Qwen3-VL-8B's 30.1
An open 8B vision-language model that reads video as a compression stream, not a stack of sampled frames.
- ElevenLabs Music v2 — New Music Model Switches Genres Mid-Track, Adds Section-Level Inpainting, and Cuts API Prices Up to 50%
A music model that swaps genres mid-track and regenerates just one section of a song.
- OpenBMB MiniCPM5-1B — 1.08B On-Device Model Reaches Open-Source SOTA in Its Size Class, Shipping GGUF and 4-Bit MLX Builds
A 1.08B on-device language model that reaches open-source SOTA in its size class.
- DeepSeek Makes Its 75% V4-Pro API Discount Permanent — Input Stays at $0.435 per Million Tokens, Output at $0.87, a Quarter of the Original Sticker Rate
DeepSeek makes its 75%-off V4-Pro API pricing permanent instead of letting the promo expire.
- Tencent Hy-MT2 — Open-Weight Multilingual Translation Family (1.8B / 7B / 30B-A3B MoE) Covers 33 Languages; 1.8B Quantizes to 440MB at 1.25-bit
An open-weight family of Tencent translation models spanning 1.8B to a 30B MoE, covering 33 languages.
- Cohere Command A+ — 218B Sparse MoE With 25B Active Params, Apache 2.0, Runs Agentic and Multimodal Workloads on Two H100s
An open-weight 218B MoE that runs agentic, multimodal, multilingual workloads on as few as two H100 GPUs.
- OpenAI Model Disproves Erdős's 1946 Unit-Distance Conjecture — First Time AI Has Autonomously Solved a Prominent Open Problem Central to a Field of Mathematics
OpenAI says an internal reasoning model autonomously disproved Paul Erdős's 1946 unit-distance conjecture, finding a new construction that beats the previously-believed best bound.
- Project Genie Plugs Into Street View — Google's World Model Generates Interactive Stylised Scenes Anchored to Real US Locations, Rolling Out to AI Ultra Subscribers Globally
Google's Genie world model now takes a real Street View address as its scene seed — pick a styled lens, walk around a generated version of that exact block.
- Stable Audio 3.0 — Stability AI Ships a Four-Model Audio Family With 6:20 Music Generation and Open Weights for Three Sizes
A four-model audio family from Stability AI with three open-weight checkpoints and full 6:20 song generation in the larger sizes.
- Qwen3.7-Max — Alibaba's New Agentic Flagship Runs 35-Hour Autonomous Tasks With 1,000+ Tool Calls
Alibaba's new flagship Qwen3.7-Max is built to run agents for tens of hours and thousands of tool calls without stalling.
- Gemini Omni — Google's 'Create Anything From Any Input' Model Ships With Conversational Video Generation
A single Gemini model that turns text, images, or audio into edited video on request.
- Gemini 3.5 Flash — Google's I/O 2026 Model Rivals Flagship Frontier Models on Coding and Agentic Tasks at 4x the Speed
Google's new Flash-tier model targets agentic coding at flagship quality and 4x the throughput.
- Cursor Composer 2.5 — Coding Model Matches Opus 4.7 and GPT-5.5 on SWE-Bench Multilingual at a Fraction of the Cost
Cursor's in-house coding model now matches frontier models on coding benchmarks at a fraction of the price.
- SANA-WM — NVIDIA's 2.6B Open-Source World Model Generates 720p One-Minute Video With 6-DoF Camera Control
An efficient open-source world model that generates minute-long, camera-controllable 720p video.
- SU-01 — Shanghai AI Lab's 31B Open-Weight Reasoner Hits Gold-Medal Scores on IMO 2025 and USAMO 2026
A 31B open-weight model that reaches gold-medal olympiad scores from a documented post-training recipe.
- Interfaze — Hybrid CNN/DNN + LLM Architecture Beats Gemini-3-Flash, Claude-Sonnet-4.6, GPT-5.4-Mini, Grok-4.3 Across 9 OCR, Vision, STT, and Structured-Output Benchmarks
A hybrid model that bolts CNNs and small task-specific networks under an LLM front-end and routes deterministic developer jobs to whichever specialized model is best.
- Cactus Compute Needle — 26M-Parameter Function-Calling Model Distilled From Gemini 3.1, MIT-Licensed, 6000 tok/s on Cactus Runtime
A 26M-parameter open model distilled from Gemini 3.1 that does nothing but call tools — small enough to run on phones, watches, and glasses.
- OpenBMB MiniCPM-V 4.6 — 1.3B Vision-Language Model Runs on iOS, Android, and HarmonyOS, Beating Qwen 3.5-0.8B at 19× Fewer Tokens
Pocket-sized vision-language model — 1.3B params, runs on phones, beats larger models per token spent.
+ 68 more in the sitemap.