New AI Model Releases — LLMs & Open Weights | Blog Network
Every new AI model worth knowing — frontier and open-weight releases, explained in plain English with the context window, parameters and benchmarks that matter.
227 releases tracked
- Apodex 1.1 — an agent model that finishes whole jobs, with 35B open weights
Apodex 1.1 works through long tasks end to end, and its 35B Mini version comes with open weights.
- Grok 4.6 on Google's agent platform — xAI's flagship arrives in Model Garden
Grok 4.6 is listed in Google's Model Garden with a 500K context window and $2 / $6 per million tokens.
- 4DAnyone — turn one handheld video of a person into a 4D model
4DAnyone reconstructs a moving person in 4D from one casual monocular video, with code and checkpoints released.
- DeepSeek V4-Flash-Vision-Exp — an experimental V4 model that reads images
An experimental DeepSeek model that adds image understanding to the V4-Flash line, at V4-Flash prices.
- Ox Alpha — an anonymous 1M-context coding model, free on OpenRouter
A free reasoning model with a 1M-token context window, live on OpenRouter with no maker attached.
- LFM2.5-DSpark — Liquid AI's draft models decode up to 3.18x faster
Three ~300M draft models that make Liquid AI's small on-device models decode two to three times faster, with identical output.
- LFM2.5 Q4_0 — Liquid AI's 4-bit models keep 96%+ of full accuracy
LFM2.5's new Q4_0 checkpoints are trained as 4-bit models rather than squeezed into 4 bits afterwards.
- Grok 4.6 on Amazon Bedrock — xAI's flagship opens to AWS teams
Grok 4.6 is generally available on Amazon Bedrock, with a US-only and a global inference profile for AWS teams.
- Ornith-1.5 — open MIT model matches Claude Opus 4.8 on Terminal-Bench
An open-weight model family that writes its own training tasks, then trains on them.
- AIDO Cell — GenBio AI's virtual cell simulates drugs on a whole human cell
A world model that holds one cell state you can perturb, clone and read out across DNA, RNA, protein and cell shape.
- GPT-5.6 Sol at half price — OpenRouter discounts OpenAI's flagship 50%
OpenAI's flagship model costs half as much through OpenRouter as it does buying from OpenAI directly.
- Toast 1 — Mixedbread's search model runs the whole retrieval loop
A dedicated search model that runs the retrieval loop itself, so the main agent stops burning tokens on it.
- Qwen3.8-27B — a 27B open model that beats Opus 4.6 Max on SWE-bench Pro
A 27B dense model with open Apache-2.0 weights that reads images and video and runs long agentic coding jobs.
- GLM-5.3 — Z.ai's coding model improves without retraining the base
Z.ai got a large jump in coding and security skill out of GLM-5.2's base model by training it harder after the fact.
- MiniMax Music 3.0 — open-weights model writes a full five-minute song
An open-weights music model that composes, arranges, performs and produces a whole song in a single pass.
- Palmyra X6 — Writer's flagship model halves the cost of an agent task
Writer's new flagship model, post-trained from GLM-5.2, runs agent jobs unattended for up to eight hours.
- Gemini 3.7 Flash — Google's coding and agent workhorse at half the price
Google's new Flash model gains 16 points on long-horizon coding work and launches at half the price of the model it replaces.
- LFM2.5-VL-3B — Liquid AI's 3B vision model reads screens on a phone
A 3.1B open-weights vision model that reads phone and desktop screens, points at objects, and calls tools on-device.
- DeepSeek raises V4 API prices — output costs more than double from August 16
DeepSeek ends its flat, ultra-cheap API rates and moves the V4 models to time-of-day pricing.
- Grok 4.6 — xAI's new flagship built for long-running agents
xAI's new flagship, improved through post-training rather than size, so it holds up over long agent runs.
- DeepSeek V4 Pro 0813 — the 1.6T flagship leaves preview
DeepSeek's 1.6-trillion-parameter flagship becomes a stable GA build after nearly four months in preview.
- Qwen3.8-2.4T-A95B — the open-weights core of Qwen3.8-Max hits Hugging Face
Qwen's 2.4T Max-class mixture-of-experts is now a public download, in BF16 and FP8.
- SL2T — Google's sign language model turns signing into text on Pixel
Google DeepMind's SL2T reads a signer's body landmarks and writes the text, and it now runs inside Gboard and Live Transcribe.
- LTX-2.5 — open-weights video model makes a 10s clip in 6.8 seconds
An open-weights video model that renders connected multi-shot scenes faster than real time.
- Nemotron 3.5 Lightning — NVIDIA's 30B open MoE for always-on agents
NVIDIA's new 30B open MoE keeps only 3B parameters active per token, aimed at agents that fire thousands of small steps.
- Motif 3 — a 314B open mixture-of-experts model under the MIT license
A 314B mixture-of-experts model with a new attention design, shipped with open MIT weights and a full technical report.
- GPT-5.6-Cyber — OpenAI splits Daybreak into Blue and Red tiers
OpenAI ships a security-specific model and puts it behind a vetted tier so approved defenders stop hitting refusals.
- Needle 2 — 14MB agentic model for phones, robots and microcontrollers
A 45M-parameter tool-calling model that ships as one 14MB binary and runs on hardware too small for anything else.
- Muse Glimmer — Meta's 30B open agentic model runs on one consumer GPU
A 30B open-weight agent model from Meta that plans, calls tools and recovers from its own errors on a single desktop GPU.
- Grok Imagine Image 2.0 — xAI's image model adds region-level editing
xAI's new image model treats editing as a first-class feature, not an add-on.
- Ling-3.0-flash — Ant Group opens 124B MoE with 5.1B active params under MIT
A 124B open-weight MoE that runs like a 5B model — Ant Group ships it under MIT.
- WeatherNext Cyclones — DeepMind's Nature-published model adds a day of warning
A neural cyclone forecaster that matches operational models one day sooner and ships open under Apache-2.0.
- Claude Fable 5 loosens biology safeguards — 85% fewer over-blocks in day-to-day use
Anthropic rewrote Fable 5's biology guardrails so ordinary health and education questions get the frontier model, not the fallback.
- ChatGPT ships smarter GPT-5.6 Sol — new reasoning slider, unlimited free chats
OpenAI retuned GPT-5.6 Sol in ChatGPT, added a reasoning slider, and made text chats unlimited for Free users.
- NVIDIA Alpamayo 2 Super — 34B open VLA model for robotaxis and self-driving
NVIDIA's largest open vision-language-action model yet, aimed at production robotaxis.
- Muse Code + Muse Spark 1.2 — Meta ships a terminal coding agent
Meta's new Muse Code coding agent lives in the terminal and ships with a co-trained Muse Spark 1.2 model.
- Maple-Preview — 20B ternary MoE trained from scratch, 218 tok/s on a Mac mini
A 20B mixture-of-experts model trained in ternary from day one, small enough to run on a laptop.
- LFM2.5-2.6B — Liquid AI's 2.6B on-device agent competes with 4x-larger models
A 2.6B open-weight agent that plans, calls tools, and runs entirely on phones, laptops, and robots.
- Shieldstral 1.0 — Mistral ships a 3B open safety classifier for text and images
A 3B open-weights safety classifier that reads a plain-language policy at inference, rates text or images, and fits in a 16GB GPU.
- Qwen3.8-Max — Alibaba's 2.4T flagship ships official with a full benchmark table
Qwen3.8-Max lands as Alibaba's official 2.4T-parameter, 95B-active MoE flagship with published benchmarks and $2/$6 per-million-token pricing.
- Seedance 2.5 — ByteDance's video model doubles to 30 seconds per generation
ByteDance's Seed lab doubles single-take video generation to 30 seconds with 50-asset multimodal referencing.
- DeepHealth Breast Ultrasound — FDA clears RadNet AI for automated lesion reads
AI that reads breast ultrasounds now has FDA sign-off, and RadNet is rolling it out across its US imaging centers.
- Inkling-Small — Thinking Machines' 276B open model matches Inkling at 1/4 the size
A 276B MoE with 12B active parameters that ships full Apache-2.0 weights and matches its 4x-larger sibling.
- Gemini Robotics ER 2 — the planning brain that watches video and coordinates robots
DeepMind's embodied-reasoning brain plans multi-step robot work, watches the video, and orchestrates multiple robots at once.
- MiniMax H3 — open-weights video model does 2K, 15s, and native stereo sound
A single open-weights video model does 2K clips, stereo audio, editing, and motion transfer up to 15 seconds long.
- DeepSeek V4-Flash goes official — 0731 hits 82.7 on Terminal Bench 2.1
DeepSeek graduates V4-Flash from preview to official with a re-post-trained checkpoint tuned for agents.
- GPT-5.6 Luna cut 80%, Terra 20% — OpenAI drops API prices three weeks after launch
OpenAI drops GPT-5.6 Luna 80% and Terra 20% three weeks after launch, passing on 20% serving-cost gains.
- Gemini Robotics 2 — DeepMind's new humanoid model controls the full body
DeepMind's new humanoid brain walks, crouches, and ties knots — and coordinates two robots at once.
- Grok Voice Think Fast 2.0 — xAI's new voice model ships at 0.70s time-to-first-audio
xAI's next-gen voice model tops the Artificial Analysis Speech-to-Speech leaderboard at sub-second latency.
- Lyria 3.5 — Google DeepMind's new music model in free Flow Music
Google DeepMind's new music model — richer melodies, clearer vocals, longer songs in the free Flow Music app.
- Microsoft Mage-VL — 4B codec-native multimodal cuts video tokens by 75%
Mage-VL treats a video like a codec — keep anchor frames, drop most predicted-frame patches, and cut visual tokens by 75%.
- KAT-Coder-V2.5-Dev — Kwaipilot's 35B open-weight agentic coding MoE
Kwaipilot opens the weights of a 35B/3B-active MoE tuned to act inside real code repositories, not just autocomplete a snippet.
- Kimi K3 open weights — Moonshot drops the 2.8T MoE on Hugging Face
The largest open-weight AI model in history — 2.8T Kimi K3 is now free to self-host.
- Microsoft MAI-Cyber-1-Flash — 96% on CyberGym at half the price of the GPT-5.4 stack
Microsoft's new security-specialist model plugs into MDASH and finds bugs cheaper and better than a GPT-5.4 stack.
- Inflect-Micro-v2 — 9.36M-parameter local text-to-speech under 40 MB
A complete English text-to-speech model that fits in 37.5 MB and beats larger systems 66% of the time in blind human tests.
- Apertus 1.5 — Switzerland's fully open 8B/70B goes multimodal
Switzerland's fully open Apertus grows eyes and ears, keeps its Apache-2.0 promise.
- FLUX 3 x mimic — the FLUX 3 backbone drives factory robots at 101 ms
A FLUX 3 spin-off that lets one on-robot GPU pilot a real factory arm in near real time.
- Claude Opus 5 — Anthropic's new Opus nears Fable 5 at half the price
Anthropic's new Opus tier comes close to Fable 5 while holding the Opus 4.8 price.
- FLUX 3 — Black Forest Labs' multimodal video, image, and robotics model
One Black Forest Labs model that generates video with audio, edits images, and drives real robots at Audi.
- Nanbeige4.2-3B — 3B looped-transformer open-weight agent hits 63.6 on SWE-Bench Verified
A 3B open-weight agent whose Looped Transformer reuses layers to match models 3–4× its size.
- Cisco Antares — 350M and 1B open-weight models that hunt code vulnerabilities
Small on-prem models that read a CWE, walk your repo, and tell you which files the vulnerability is hiding in.
- Microsoft Mage-Flow — 4B native-resolution image model that keeps up with 20B systems
A 4B image model from Microsoft Research that generates or edits at any resolution and matches open systems five times its size.
- Solar Open 2 — Upstage's 250B/15B open-weight MoE built for agentic use
Korea's Upstage ships a 250B/15B open-weight MoE with 1M context and a hybrid-attention stack, aimed at agentic work.
- Poolside Laguna S 2.1 — 118B open-weight coding MoE with 8B active
Poolside ships Laguna S 2.1 — a 118B/8B-active open-weight coding MoE with a 1M-token context, pitched as the West's answer to DeepSeek and Qwen.
- Gemini 3.6 Flash — Google's workhorse plus Flash-Lite and Flash Cyber
Google ships three Flash-tier Gemini variants tuned for coding, high-throughput agents, and cybersecurity.
- Qwen-Image-3.0 — Alibaba's third-gen image model ships without weights
Alibaba's third-gen image model takes 4,500-token prompts and renders dense text — but ships without weights, license, or benchmarks.
- Qwen-Audio-3.0-TTS — Alibaba's TTS hits #1 on Artificial Analysis
Alibaba's new hosted TTS ships in Flash and Plus tiers, spans 16 languages, and takes #1 on the Artificial Analysis speech leaderboard.
- Qwen3.8-Max Preview — Alibaba's 2.4T multimodal flagship arrives, weights promised
Alibaba announces Qwen3.8-Max, a 2.4T multimodal preview, and promises open weights soon.
- NVIDIA Cosmos 3 Edge — 4B world model that runs physical AI on Jetson
NVIDIA's Cosmos family gets a small, on-device sibling built for real-time robotics.
- OvisOCR2 — 0.8B Alibaba model tops OmniDocBench and beats pipeline OCR
Alibaba's 0.8B end-to-end document parser sets state of the art on OmniDocBench v1.6.
- Nemotron 3 Embed — NVIDIA's open 8B embedder takes #1 on RTEB
NVIDIA ships an open 8B embedder that takes #1 on RTEB, plus two efficient 1B variants for production RAG.
- Kimi K3 — Moonshot's 2.8T flagship with 1M context lands on web, app, and API
Moonshot AI ships Kimi K3 today, jumping to a 2.8T mixture-of-experts with a 1M-token context and native vision.
- Descartes — Hemispheric's frontier NeuroAI model for decoding EEG brain signals
A 6B-parameter AI model that converts 15-minute EEG recordings into quantitative brain-health diagnostics.
- NeuroVFM — brain-scan AI outperforms GPT-5 on clinical triage
Michigan Medicine's 5M-scan neuroimaging AI trained on uncurated hospital data outperforms GPT-5 on triage at 24× lower cost.
- Boogu-Image-0.1 — 10B open image model trained on 208M images for ~$400K
Apache-2.0 10B image generation model reports near closed-source quality after training on 208M images for about $400K in compute.
- MonkeyOCRv2 — visual-text foundation model for document AI
MonkeyOCRv2 pretrains a visual-text foundation model on 113M multilingual document images.
- Inkling — Thinking Machines' first open-weights 975B/41B multimodal MoE
Mira Murati's Thinking Machines Lab ships its first foundation model, and puts the full weights on HuggingFace under Apache 2.0.
- Xiaomi-Robotics-U0 — open 38B unified embodied world model
38B open autoregressive foundation model that generates images, embodied scenes and robot video from one shared tokenizer.
- Bonsai 27B — first 27B-class LLM to run on a phone at ~4 GB
Bonsai 27B is the first 27B-class open-weight model small enough to run on a phone — 3.9 GB of weights, 262K context, Apache-2.0.
- Seedream 5.0 Pro — ByteDance's image model reasons before it draws
ByteDance's Seedream 5.0 Pro thinks about the prompt before it draws, then hits 2K with clean text in 14 languages.
- GPT-5.6 goes public — Sol, Terra, and Luna clear the White House gate
OpenAI's frontier model line goes broad after the government finishes its second look.
- Muse Spark 1.1 — Meta MSL opens its first paid API at $1.25 / $4.25 per million tokens
Meta finally opens Muse Spark to outside developers — with a price tag under Claude and GPT.
- Grok 4.5 — xAI's Opus-class flagship at $2/$6 per million tokens
xAI's new coding flagship ties GPT-5.5 on Terminal-Bench at a quarter of the price of Fable 5 and Opus.
- OpenAI GPT-Live — full-duplex voice model now powers ChatGPT Voice
OpenAI GPT-Live is a full-duplex voice model that lets ChatGPT listen and speak at the same time.
- SWE-1.7 — Cognition's coding model runs on Devin at 1000 tok/s via Cerebras
SWE-1.7 is Cognition's new coding model — near GPT-5.5 on agentic coding, running at 1000 tokens per second on Devin.
- Robostral Navigate — Mistral's first embodied model, 8B, single RGB camera
Mistral's first embodied model steers wheeled, legged, and flying robots from a single RGB camera and a plain-language instruction.
- Claude Fable 5 access extended again — Anthropic pushes subscription window to July 19
Anthropic gives Fable 5 subscribers another week before usage credits kick in.
- Cohere Transcribe Arabic — 2B open-weight ASR beats Whisper on dialect and code-switching
The open-source Arabic speech model that finally beats Whisper on dialect audio.
- Meta Muse Image — MSL's first image model plans, calls tools, and self-refines
Meta Superintelligence Labs' first image model plans, calls tools, and self-refines like a reasoning LLM.
- Ternlight — 7 MB embedding model that runs in the browser
A 7 MB sentence-embedding model that runs in the browser at ~5 ms per query, with no API call.
- xAI Grok Voice — 21 new flagship voices with speech tags for pacing
Grok Voice now ships 21 new multilingual flagship voices plus inline speech tags for pacing and whispering.
- OpenAI gpt-realtime-2.1 — voice model gains reasoning-effort dial and a mini variant
gpt-realtime-2.1 adds a reasoning-effort dial and better silence/interruption handling to OpenAI's Realtime API.
- Tencent Hunyuan Hy3 — 295B/21B MoE open-sourced with 256K context
Tencent open-sources its Hunyuan flagship: 295B total, 21B active, 256K context, Apache 2.0.
- Leanstral 1.5 — Mistral's updated Lean 4 formal-proof model
Mistral's Lean 4 theorem-prover saturates miniF2F, tops PutnamBench, and ships as an Apache-2.0 119B/6.5B MoE.
- TabFM — Google's zero-shot foundation model for tabular data
A pretrained-once foundation model that skips XGBoost's tuning ritual for tabular classification and regression.
- Gemini Omni Flash + Nano Banana 2 Lite — Google's new video and image models
Google ships two preview models: a 10-second video generator priced per second, and a 4-second image generator priced per shot.
- Claude Sonnet 5 — Anthropic's new agentic Sonnet at Opus-class quality
Anthropic's new mid-tier Sonnet 5 lands with a 1M-token context, adaptive thinking, and $3/$15 pricing — Opus-class quality without Opus pricing.
- Agents-A1 — Shanghai AI Lab 35B MoE matches trillion-parameter agents
Open-weight 35B agent from Shanghai AI Lab posts SOTA on SEAL-0, IFBench, and FrontierScience-Research.
- LongCat-2.0 — Meituan's 1.6T open-source MoE for agentic coding
Meituan's first frontier-tier open-source coding model — 1.6T parameters with 48B active, trained without a single NVIDIA chip.
- GPT-5.6 — OpenAI previews Sol, Terra, and Luna tiers
OpenAI's new generation splits into three named tiers and adds an ultra mode that wires subagents into the flagship model.
- Ornith 1.0 — open-weight coding models that learn their own RL scaffold
Open-weight agentic coding models that learn to write their own RL scaffold instead of relying on a fixed harness.
- Gemini 3.5 Flash gets Computer Use — native browser, mobile, and desktop agents
Gemini 3.5 Flash can now see a screen and click, type, and scroll on its own through a single built-in tool.
- Krea 2 — open-weight 12B image model with 2-second Turbo variant
Krea AI open-sources a 12B Diffusion Transformer image model with a Turbo variant that draws 2K in two seconds.
- Mistral OCR 4 — 170-language document model with bounding boxes and confidence scores
Mistral's new document model returns structured pages with boxes, block types, and per-word confidence at $4 per 1,000.
- Baidu Unlimited-OCR — 3B vision model parses long documents in one pass
Baidu's open 3B OCR model swaps standard attention for R-SWA so it can transcribe dozens of pages without the usual KV-cache blowup.
- PP-OCRv6 — PaddlePaddle ships 50-language OCR family from 1.5M to 34.5M params
PaddlePaddle's PP-OCRv6 is a three-tier OCR family — Tiny 1.5M to Medium 34.5M — that recognises 50 languages and beats PP-OCRv5_server.
- Sakana Fugu — multi-agent orchestration model that matches Fable 5 on quality
Sakana AI ships Fugu, a single API that routes each request to a pool of frontier models and verifies the answer before returning it.
- MolmoMotion — Ai2's language-guided 3D motion forecasting models
MolmoMotion predicts how points on objects move in 3D from a video frame and a text instruction, with weights, a 1.16M-video dataset, and a benchmark.
- Qwen-Robot Suite — Alibaba's three foundation models for robots
Three open foundation models from Alibaba's Qwen team that move robots, drive them around, and predict what happens next.
- VibeThinker-3B — Weibo's 3B reasoning model hits 80.2% on LiveCodeBench v6
Sina Weibo's 3B model finetuned from Qwen2.5-Coder-3B, MIT-licensed, scoring 94.3 on AIME26 and 80.2 on LiveCodeBench v6.
- Grok Imagine Video 1.5 — xAI's image-to-video model goes GA at $0.14/sec 720p
xAI's image-to-video model — the engine behind Grok Imagine's video clips — is now generally available as a pay-per-second API.
- GLM-5.2 — Z.ai's new flagship coding model with 1M context
Z.ai's new coding flagship lands first inside the GLM Coding Plan, with API, chatbot, and open weights set for next week.
- Kimi K2.7-Code — Moonshot's 1T MoE Coding Model Beats Claude Opus on MCPMark
Moonshot's coding-specialized fork of Kimi K2.6, with faster reasoning and a higher MCP tool-use score than Claude Opus 4.8.
- Decart Ships Oasis 3 — First API-Accessible Interactive World Model for Physical AI Streams Three Synchronized 768×512 Camera Views at 22 FPS With <200ms End-to-End Latency on NVIDIA HGX B200, Priced at $0.02 per Second of Simulation
Decart's Oasis 3 lets robotics and AV teams stream three synchronized photorealistic camera views in real time from a text prompt, action-conditioned, via API.
- Amap Open-Sources ABot-Earth 0.5 — Alibaba's 3D Native World Model Generates Kilometer-Scale 3D Gaussian-Splatting City Scenes From a Single Satellite Image or Text Prompt in About 10 Minutes per km² on a Consumer GPU
Alibaba's Amap turns one satellite image into an interactive 3D city in roughly 10 minutes per square kilometer.
- Google Ships DiffusionGemma — Apache-2.0 26B/3.8B-Active Mixture-of-Experts That Denoises 256 Tokens in Parallel via Discrete Block Diffusion, Hits 1,000+ Tokens/Sec on H100 and 700+ on RTX 5090 While Posting 77.6% MMLU Pro, 73.2% GPQA Diamond, and 69.1% LiveCodeBench v6
Google opens a 26B-parameter Gemma that denoises 256 tokens at once instead of generating them one by one.
- Kuaishou Open-Sources Keye-VL-2.0-30B-A3B — Apache-2.0 Mixture-of-Experts Vision-Language Model Activates 3B Parameters Per Token, Lands 256K Context With DeepSeek Sparse Attention for Lossless Long-Video Reasoning, Beats Qwen3-VL-235B on LongVideoBench at 74.1, and Tops 70.1 mIoU on QVHighlights-TimeLens
A 30B/3B-active MoE that pushes long-video understanding past dense 235B baselines, all under Apache-2.0.
- Cohere Ships North Mini Code 1.0 — Apache-2.0 Sparse 30B/3B-Active Mixture-of-Experts Coding Model With a 256K Context, 64K Max Output, 83.2% Pass@1 on SWE-Bench Verified, 63% on Terminal-Bench v2, and 2.8× the Output Throughput of Devstral Small 2
Cohere's first model aimed squarely at developers — a 30B/3B-active MoE coding model under Apache 2.0.
- Anthropic Ships Claude Fable 5 and Claude Mythos 5 — New Flagship Tops Hebbia's Finance Benchmark and Cognition's FrontierCode Eval at $10/$50 per Million Tokens, With Mythos 5 Restricted to Project Glasswing Partners and Select Biology Researchers
Fable 5 ships to everyone today; Mythos 5 stays gated to Glasswing partners and biology researchers.
- Google DeepMind Ships Gemini 3.5 Live Translate — Audio Model Streams Near Real-Time Speech-to-Speech in 70+ Languages With Preserved Intonation, SynthID Watermarks, and Public Preview on Gemini Live API, AI Studio, Meet, and Google Translate
First production speech-to-speech model that translates continuously instead of taking turns.
+ 107 more in the sitemap.