New AI Model Releases — LLMs & Open Weights
Every new AI model worth knowing — frontier and open-weight releases, explained in plain English with the context window, parameters and benchmarks that matter.
276 releases tracked
- Qwen3.8-Omni-Flash — Alibaba's omni model gets a 1M-token context
Alibaba's omnimodal Flash model takes audio and video in a 1M-token context, at a fraction of the old audio price.
- Bonsai 2 27B — ternary model keeps 98.2% of full precision in 5.9 GB
A 27B model squeezed into ternary weights that gives up 1.5 points of benchmark score for a 9x smaller download.
- LimiX-2 — a 400M model tops three structured-data benchmarks
One 400M model handles classification, regression and missing values on tables, with no per-dataset training.
- Jev — TypeSafe's model returns typed decisions, not text
A model that answers typed questions with calibrated probabilities instead of writing prose you have to parse.
- Gemini 3.8 Live — Google's voice models talk while they think
Google's new Live API models hold a real-time voice conversation and keep reasoning in the background while they speak.
- Salesforce Koa — a CRM reasoning model built on NVIDIA Nemotron
Salesforce Koa is a CRM-focused reasoning model that Salesforce post-trained from NVIDIA's open-weight Nemotron 3 Super.
- YuE2-3B — open music model tops Suno v5 and v6 on WildSongBench
An open 3B music model that writes an editable score first, then renders a full song with vocals.
- Fugu Max and Fugu Ultra v2 — Sakana's router splits into cheap and strong
Sakana AI splits its Fugu orchestrator into a cheap tier and a high-capability tier, both behind one OpenAI-compatible API.
- GPT-Live-1 in the API — full-duplex voice for $0.05 a minute
OpenAI's full-duplex voice model reaches the API, so any app can listen and talk at the same time.
- North Small Translate — Cohere's open translation model beats DeepL on WMT26
Cohere Labs opens the weights of a 218B mixture-of-experts model built only for translation across 50 languages.
- SWE-2 — Cognition's coding model lands within a point of Fable 5.1
SWE-2 post-trains Moonshot's Kimi K3 with reinforcement learning and lands near frontier coding scores for a fraction of the price.
- Nex-N2.5 — three open-weight agent models, up to 1.6 trillion parameters
Nex AGI releases mini, Pro and Max — agent models that treat vision as a working interface, not just an input.
- Suno v6 — a music model family trained only on licensed catalogues
Suno v6 is the company's first model generation built with record labels as licensing partners rather than around them.
- Gander — an open 9B model that listens, watches and works at once
Gander pairs a fast streaming speech model with a slower reasoning agent, so a voice conversation keeps flowing while long tasks run.
- DeepSeek V4.1 Flash — a 552B open-weight rebuild with 1M context
DeepSeek's smallest new-architecture model is now generally available with MIT open weights, a 1M-token context and top agentic scores.
- AuK — Tencent's open speech model generates and edits audio by instruction
One 1.5B open model that writes, rewrites, cleans and restyles speech from a sentence of instructions.
- NeoHorse-1 — open 4B and 9B models post-trained by a routing harness
TokenRhythm routes agent tasks across a pool of models, keeps what worked, and trains 4B and 9B models on it.
- Mercury 2.5 — Inception's diffusion model hits 1,107 tokens per second
The largest diffusion language model trained so far, streaming over 1,100 tokens a second.
- ChatGPT Images 2.5 — OpenAI's image model adds Sketch and two API tiers
OpenAI's Images 2.5 keeps the subject of a reference photo intact while you edit everything around it.
- MiniCPM5-2B — a 2B open model that leads the sub-4B field
OpenBMB's new 2B model takes the top spot among open models under 4B parameters.
- Lyria 3.5 comes to Gemini — Google's music model lands in the app and API
Google's music model is now one API call away, and free for every Gemini user.
- MAI-Image-2.6 — Microsoft's image model lands at No. 2 on Arena
Microsoft's newest image model ranks second on Arena and ships with a faster, cheaper Flash variant.
- Puffin-World — an open 3D world model with physics, depth and camera
One open model reads camera pose, depth and physics from images, then generates and reconstructs 3D scenes.
- OpenEvidence model family — Osler, Sackett and Snow ship free to clinicians
Four medical AI models that all aim at the same accuracy — what changes is how long each one thinks.
- humain-m3 — a 428B Arabic model HUMAIN commissioned from MiniMax
Saudi Arabia's HUMAIN commissioned a 428B Arabic frontier model from MiniMax and opened it as a limited preview.
- LLaDA-Image — a 6B open image generator with a 4-step turbo variant
A 6B diffusion model from Ant Group that writes images from text, edits them from a reference, and publishes its training recipe.
- GPT-6 Astra — OpenAI's computer-use model starts rolling out
OpenAI's new frontier model drives software like a person does, and it is the first to hit the company's Critical cyber bar.
- K2 Horizon — six fully open models, from 0.9B to 375B
Six Apache-2.0 models that share one recipe, published with the training data and code behind them.
- WeatherNext 3 — DeepMind's weather model goes hourly at 5km
DeepMind's global forecast model moves from 6-hour, 25km predictions to hourly ones at 5km.
- Quasar 438B — Multiverse Computing's first large model, built in Europe
Multiverse Computing's first large model: 438B parameters, aimed at enterprise agents and coding.
- Muse Spark 1.3 — Meta's flagship model uses 25% fewer tokens on coding
Meta's Muse Spark 1.3 keeps the same price and 1M-token window while doing coding work with fewer tool calls and fewer tokens.
- Gemini 3.8 Flash — Google's best coding model yet, plus a cyber twin
Google's new Flash model keeps the 3.7 price while scoring higher on coding, reasoning and security work.
- Qwen3.8-Max-0902 — Alibaba's 2.4T flagship gets a coding refresh
A dated snapshot of Qwen3.8-Max, post-trained for bigger codebases and longer unsupervised agent runs.
- Atlas — World Labs' omni model for text, image, video and 3D
World Labs' Atlas is one model for text, images, video and 3D, with pixel-level control of the camera.
- Claude Fable 5.1 — Anthropic's new top model, with cache reads 75% cheaper
Anthropic's new top-end Claude keeps the same $10/$50 rates but cuts cache reads by 75%, and jumps on long agent tasks.
- TimesFM-3 — Google's forecasting model handles many series at once
Google Research's 330M forecasting model predicts many linked time series in a single pass, with no fine-tuning.
- DeepSeek-V4-Flash-Vision-Exp weights go public — 305B multimodal MoE under MIT
DeepSeek's first multimodal V4 model leaves API-only preview — 305B weights, MIT license, and inference code on Hugging Face.
- H3 Max — fal's post-trained MiniMax H3 makes a 5-second clip in under 3 seconds
A post-trained MiniMax H3 that keeps the quality and returns about 35 times more clips per second.
- GLM-5.3 weights go public — Z.ai's 753B coding model lands on Hugging Face
Z.ai's 753B GLM-5.3 coding model is now a public download on Hugging Face, in BF16 and FP8.
- Tencent Hy4 preview — 770B open-weights model with a 1M-token context
Tencent open-sources Hy4 preview, a 770B mixture-of-experts model with 49B active parameters and a 1M-token context.
- Cohere Parse — a 2.3B document model that turns PDFs into clean Markdown
A small vision model built for document pipelines: Markdown out, tables kept, 4.5 pages a second, $1.50 per thousand pages.
- Gemini Omni 1.1 Flash — Google's video model gets keyframes and 4K output
Google's video model now takes direction: fixed keyframes, 40-second scenes, cheap 360p drafts and 4K finals.
- Gemini 3.5 Transcribe — Google's speech model cleans up your ums and ahs
Google's new speech-to-text model writes the sentence you meant, not the one you stumbled through.
- GigaBrain-0.7 — an open robot brain trained on 37,000 hours of embodied data
GigaBrain-0.7 is an Apache-2.0 vision-language-action model that turns camera views and a spoken instruction into robot movement.
- Qwen3.8-Flash-Next — open 125B MoE with 6B active previews Qwen4
Qwen3.8-Flash-Next ships open weights for the first public preview of the Qwen4 architecture.
- WeMM-Embedding — Tencent's multimodal retrieval models top MMEB-v2
Tencent open-sourced WeMM-Embedding, three multimodal embedding models that turn text, images, video and documents into one vector space.
- IBM Granite 4.2 — open reasoning models with a thinking switch
IBM's first dense reasoning models, in 3B, 8B and 30B, with Apache-2.0 weights and a thinking mode you can switch off.
- GPT-5.6 in Kiro — OpenAI's model family lands in AWS's coding agent
OpenAI's GPT-5.6 family is now selectable inside Kiro, AWS's spec-driven coding agent, for the first time.
- Thomson 1.0 Small — Thomson Reuters ships its own 35B legal and tax model
A legal and tax specialist model, trained on Westlaw-grade content, with a small open-weight version anyone can download.
- Apodex 1.1 — an agent model that finishes whole jobs, with 35B open weights
Apodex 1.1 works through long tasks end to end, and its 35B Mini version comes with open weights.
- Grok 4.6 on Google's agent platform — xAI's flagship arrives in Model Garden
Grok 4.6 is listed in Google's Model Garden with a 500K context window and $2 / $6 per million tokens.
- 4DAnyone — turn one handheld video of a person into a 4D model
4DAnyone reconstructs a moving person in 4D from one casual monocular video, with code and checkpoints released.
- DeepSeek V4-Flash-Vision-Exp — an experimental V4 model that reads images
An experimental DeepSeek model that adds image understanding to the V4-Flash line, at V4-Flash prices.
- GLM-5.3-Flash — Z.ai opens the weights of the model that was Ox Alpha
The stealth model that topped OpenRouter now has a name, a model card and MIT-licensed weights: GLM-5.3-Flash.
- LFM2.5-DSpark — Liquid AI's draft models decode up to 3.18x faster
Three ~300M draft models that make Liquid AI's small on-device models decode two to three times faster, with identical output.
- LFM2.5 Q4_0 — Liquid AI's 4-bit models keep 96%+ of full accuracy
LFM2.5's new Q4_0 checkpoints are trained as 4-bit models rather than squeezed into 4 bits afterwards.
- Grok 4.6 on Amazon Bedrock — xAI's flagship opens to AWS teams
Grok 4.6 is generally available on Amazon Bedrock, with a US-only and a global inference profile for AWS teams.
- Ornith-1.5 — open MIT model matches Claude Opus 4.8 on Terminal-Bench
An open-weight model family that writes its own training tasks, then trains on them.
- AIDO Cell — GenBio AI's virtual cell simulates drugs on a whole human cell
A world model that holds one cell state you can perturb, clone and read out across DNA, RNA, protein and cell shape.
- GPT-5.6 Sol at half price — OpenRouter discounts OpenAI's flagship 50%
OpenAI's flagship model costs half as much through OpenRouter as it does buying from OpenAI directly.
- Toast 1 — Mixedbread's search model runs the whole retrieval loop
A dedicated search model that runs the retrieval loop itself, so the main agent stops burning tokens on it.
- Qwen3.8-27B — a 27B open model that beats Opus 4.6 Max on SWE-bench Pro
A 27B dense model with open Apache-2.0 weights that reads images and video and runs long agentic coding jobs.
- GLM-5.3 — Z.ai's coding model improves without retraining the base
Z.ai got a large jump in coding and security skill out of GLM-5.2's base model by training it harder after the fact.
- MiniMax Music 3.0 — open-weights model writes a full five-minute song
An open-weights music model that composes, arranges, performs and produces a whole song in a single pass.
- Palmyra X6 — Writer's flagship model halves the cost of an agent task
Writer's new flagship model, post-trained from GLM-5.2, runs agent jobs unattended for up to eight hours.
- Gemini 3.7 Flash — Google's coding and agent workhorse at half the price
Google's new Flash model gains 16 points on long-horizon coding work and launches at half the price of the model it replaces.
- LFM2.5-VL-3B — Liquid AI's 3B vision model reads screens on a phone
A 3.1B open-weights vision model that reads phone and desktop screens, points at objects, and calls tools on-device.
- DeepSeek raises V4 API prices — output costs more than double from August 16
DeepSeek ends its flat, ultra-cheap API rates and moves the V4 models to time-of-day pricing.
- Grok 4.6 — xAI's new flagship built for long-running agents
xAI's new flagship, improved through post-training rather than size, so it holds up over long agent runs.
- DeepSeek V4 Pro 0813 — the 1.6T flagship leaves preview
DeepSeek's 1.6-trillion-parameter flagship becomes a stable GA build after nearly four months in preview.
- Qwen3.8-2.4T-A95B — the open-weights core of Qwen3.8-Max hits Hugging Face
Qwen's 2.4T Max-class mixture-of-experts is now a public download, in BF16 and FP8.
- SL2T — Google's sign language model turns signing into text on Pixel
Google DeepMind's SL2T reads a signer's body landmarks and writes the text, and it now runs inside Gboard and Live Transcribe.
- LTX-2.5 — open-weights video model makes a 10s clip in 6.8 seconds
An open-weights video model that renders connected multi-shot scenes faster than real time.
- Nemotron 3.5 Lightning — NVIDIA's 30B open MoE for always-on agents
NVIDIA's new 30B open MoE keeps only 3B parameters active per token, aimed at agents that fire thousands of small steps.
- Motif 3 — a 314B open mixture-of-experts model under the MIT license
A 314B mixture-of-experts model with a new attention design, shipped with open MIT weights and a full technical report.
- GPT-5.6-Cyber — OpenAI splits Daybreak into Blue and Red tiers
OpenAI ships a security-specific model and puts it behind a vetted tier so approved defenders stop hitting refusals.
- Needle 2 — 14MB agentic model for phones, robots and microcontrollers
A 45M-parameter tool-calling model that ships as one 14MB binary and runs on hardware too small for anything else.
- Muse Glimmer — Meta's 30B open agentic model runs on one consumer GPU
A 30B open-weight agent model from Meta that plans, calls tools and recovers from its own errors on a single desktop GPU.
- Grok Imagine Image 2.0 — xAI's image model adds region-level editing
xAI's new image model treats editing as a first-class feature, not an add-on.
- Ling-3.0-flash — Ant Group opens 124B MoE with 5.1B active params under MIT
A 124B open-weight MoE that runs like a 5B model — Ant Group ships it under MIT.
- WeatherNext Cyclones — DeepMind's Nature-published model adds a day of warning
A neural cyclone forecaster that matches operational models one day sooner and ships open under Apache-2.0.
- Claude Fable 5 loosens biology safeguards — 85% fewer over-blocks in day-to-day use
Anthropic rewrote Fable 5's biology guardrails so ordinary health and education questions get the frontier model, not the fallback.
- ChatGPT ships smarter GPT-5.6 Sol — new reasoning slider, unlimited free chats
OpenAI retuned GPT-5.6 Sol in ChatGPT, added a reasoning slider, and made text chats unlimited for Free users.
- NVIDIA Alpamayo 2 Super — 34B open VLA model for robotaxis and self-driving
NVIDIA's largest open vision-language-action model yet, aimed at production robotaxis.
- Muse Code + Muse Spark 1.2 — Meta ships a terminal coding agent
Meta's new Muse Code coding agent lives in the terminal and ships with a co-trained Muse Spark 1.2 model.
- Maple-Preview — 20B ternary MoE trained from scratch, 218 tok/s on a Mac mini
A 20B mixture-of-experts model trained in ternary from day one, small enough to run on a laptop.
- LFM2.5-2.6B — Liquid AI's 2.6B on-device agent competes with 4x-larger models
A 2.6B open-weight agent that plans, calls tools, and runs entirely on phones, laptops, and robots.
- Shieldstral 1.0 — Mistral ships a 3B open safety classifier for text and images
A 3B open-weights safety classifier that reads a plain-language policy at inference, rates text or images, and fits in a 16GB GPU.
- Qwen3.8-Max — Alibaba's 2.4T flagship ships official with a full benchmark table
Qwen3.8-Max lands as Alibaba's official 2.4T-parameter, 95B-active MoE flagship with published benchmarks and $2/$6 per-million-token pricing.
- Seedance 2.5 — ByteDance's video model doubles to 30 seconds per generation
ByteDance's Seed lab doubles single-take video generation to 30 seconds with 50-asset multimodal referencing.
- DeepHealth Breast Ultrasound — FDA clears RadNet AI for automated lesion reads
AI that reads breast ultrasounds now has FDA sign-off, and RadNet is rolling it out across its US imaging centers.
- Inkling-Small — Thinking Machines' 276B open model matches Inkling at 1/4 the size
A 276B MoE with 12B active parameters that ships full Apache-2.0 weights and matches its 4x-larger sibling.
- Gemini Robotics ER 2 — the planning brain that watches video and coordinates robots
DeepMind's embodied-reasoning brain plans multi-step robot work, watches the video, and orchestrates multiple robots at once.
- MiniMax H3 — open-weights video model does 2K, 15s, and native stereo sound
A single open-weights video model does 2K clips, stereo audio, editing, and motion transfer up to 15 seconds long.
- DeepSeek V4-Flash goes official — 0731 hits 82.7 on Terminal Bench 2.1
DeepSeek graduates V4-Flash from preview to official with a re-post-trained checkpoint tuned for agents.
- GPT-5.6 Luna cut 80%, Terra 20% — OpenAI drops API prices three weeks after launch
OpenAI drops GPT-5.6 Luna 80% and Terra 20% three weeks after launch, passing on 20% serving-cost gains.
- Gemini Robotics 2 — DeepMind's new humanoid model controls the full body
DeepMind's new humanoid brain walks, crouches, and ties knots — and coordinates two robots at once.
- Grok Voice Think Fast 2.0 — xAI's new voice model ships at 0.70s time-to-first-audio
xAI's next-gen voice model tops the Artificial Analysis Speech-to-Speech leaderboard at sub-second latency.
- Lyria 3.5 — Google DeepMind's new music model in free Flow Music
Google DeepMind's new music model — richer melodies, clearer vocals, longer songs in the free Flow Music app.
- Microsoft Mage-VL — 4B codec-native multimodal cuts video tokens by 75%
Mage-VL treats a video like a codec — keep anchor frames, drop most predicted-frame patches, and cut visual tokens by 75%.
- KAT-Coder-V2.5-Dev — Kwaipilot's 35B open-weight agentic coding MoE
Kwaipilot opens the weights of a 35B/3B-active MoE tuned to act inside real code repositories, not just autocomplete a snippet.
- Kimi K3 open weights — Moonshot drops the 2.8T MoE on Hugging Face
The largest open-weight AI model in history — 2.8T Kimi K3 is now free to self-host.
- Microsoft MAI-Cyber-1-Flash — 96% on CyberGym at half the price of the GPT-5.4 stack
Microsoft's new security-specialist model plugs into MDASH and finds bugs cheaper and better than a GPT-5.4 stack.
- Inflect-Micro-v2 — 9.36M-parameter local text-to-speech under 40 MB
A complete English text-to-speech model that fits in 37.5 MB and beats larger systems 66% of the time in blind human tests.
- Apertus 1.5 — Switzerland's fully open 8B/70B goes multimodal
Switzerland's fully open Apertus grows eyes and ears, keeps its Apache-2.0 promise.
- FLUX 3 x mimic — the FLUX 3 backbone drives factory robots at 101 ms
A FLUX 3 spin-off that lets one on-robot GPU pilot a real factory arm in near real time.
- Claude Opus 5 — Anthropic's new Opus nears Fable 5 at half the price
Anthropic's new Opus tier comes close to Fable 5 while holding the Opus 4.8 price.
- FLUX 3 — Black Forest Labs' multimodal video, image, and robotics model
One Black Forest Labs model that generates video with audio, edits images, and drives real robots at Audi.
- Nanbeige4.2-3B — 3B looped-transformer open-weight agent hits 63.6 on SWE-Bench Verified
A 3B open-weight agent whose Looped Transformer reuses layers to match models 3–4× its size.
- Cisco Antares — 350M and 1B open-weight models that hunt code vulnerabilities
Small on-prem models that read a CWE, walk your repo, and tell you which files the vulnerability is hiding in.
- Microsoft Mage-Flow — 4B native-resolution image model that keeps up with 20B systems
A 4B image model from Microsoft Research that generates or edits at any resolution and matches open systems five times its size.
- Solar Open 2 — Upstage's 250B/15B open-weight MoE built for agentic use
Korea's Upstage ships a 250B/15B open-weight MoE with 1M context and a hybrid-attention stack, aimed at agentic work.
- Poolside Laguna S 2.1 — 118B open-weight coding MoE with 8B active
Poolside ships Laguna S 2.1 — a 118B/8B-active open-weight coding MoE with a 1M-token context, pitched as the West's answer to DeepSeek and Qwen.
- Gemini 3.6 Flash — Google's workhorse plus Flash-Lite and Flash Cyber
Google ships three Flash-tier Gemini variants tuned for coding, high-throughput agents, and cybersecurity.
- Qwen-Image-3.0 — Alibaba's third-gen image model ships without weights
Alibaba's third-gen image model takes 4,500-token prompts and renders dense text — but ships without weights, license, or benchmarks.
- Qwen-Audio-3.0-TTS — Alibaba's TTS hits #1 on Artificial Analysis
Alibaba's new hosted TTS ships in Flash and Plus tiers, spans 16 languages, and takes #1 on the Artificial Analysis speech leaderboard.
- Qwen3.8-Max Preview — Alibaba's 2.4T multimodal flagship arrives, weights promised
Alibaba announces Qwen3.8-Max, a 2.4T multimodal preview, and promises open weights soon.
- NVIDIA Cosmos 3 Edge — 4B world model that runs physical AI on Jetson
NVIDIA's Cosmos family gets a small, on-device sibling built for real-time robotics.
- OvisOCR2 — 0.8B Alibaba model tops OmniDocBench and beats pipeline OCR
Alibaba's 0.8B end-to-end document parser sets state of the art on OmniDocBench v1.6.
- Nemotron 3 Embed — NVIDIA's open 8B embedder takes #1 on RTEB
NVIDIA ships an open 8B embedder that takes #1 on RTEB, plus two efficient 1B variants for production RAG.
+ 156 more in the sitemap.