New AI Tools — Products, CLIs & IDE Features | Blog Network
The newest AI tools, products, CLIs and IDE features — what shipped, who it's for, and how to try it tonight, in plain English.
478 releases tracked
- Codex CLI 0.162.0 — managed Git worktrees and pinned tasks
Codex CLI can now create and list Git worktrees itself and lets you pin tasks in its Command Center.
- bigarrow — AI agents point at what to click with big arrows on your Mac
A finger for your coding agent: it points at the button you need to press, then disappears.
- Gemini agent — Google Cloud's one agent for work runs tasks for days
Google Cloud's Gemini agent takes a goal, not step-by-step instructions, and comes back with finished work.
- Claude Code 2.1.295 — hooks can fail closed, terminals show agent status
Claude Code lets a broken safety hook block the action, and tells your terminal what the agent is doing.
- Codex CLI 0.161.0 — GPT-6.1 Sol becomes the default model
Codex CLI switches its default model to GPT-6.1 Sol and signs in to MCP servers from the terminal.
- Microsoft Execution Containers (MXC) — Windows agent sandbox goes GA
Windows now gives AI agents an OS-enforced sandbox, and Microsoft is pushing more models to run locally on PCs.
- Google Playground — Google Labs turns text prompts into playable browser games
Describe a game in plain words, play it straight away in the browser, then share it with a link.
- Docker Agent 1.149 — Docker's YAML agent runtime loads skills from GitHub
Write an agent as a YAML file, give it MCP tools, and run or share it like a container.
- ChatGPT Intelligent UI — GPT-6 answers with charts, buttons and mini-apps
ChatGPT answers can now be small interactive tools instead of only text.
- e2e 0.18 — TesterArmy's AI testing framework tops GitHub trending
Write the goal, not the selectors: e2e lets an agent click through your app, then caches what worked.
- REA 4.1 — the agent reverse-engineering kit now reads Android APKs and firmware
Point your coding agent at an app or binary and it can decompile it, trace a feature and show the evidence.
- Claude Code 2.1.290 — WebFetch reads past 100K characters, /loop survives compaction
Claude Code stops losing long web pages and scheduled tasks, and gives mods more to read.
- Strata v0.1.39 — Qwen3.8-Flash-Next 125B on a 12 GB gaming GPU
A one-click engine that splits a 125B open model across your GPU, RAM, CPU and SSD so it runs on a normal gaming PC.
- Claude Code 2.1.289 — deny rules now catch env-prefixed shell commands
Claude Code tightens its shell permission rules and makes plugins and mods fail alone instead of ending the session.
- Muse Gadgets — Meta open-sources SDKs to build hardware for Muse
Meta's open-source kit turns ESP32 boards and Raspberry Pis into devices its Muse agent can see, hear and control.
- Claude Code 2.1.288 — timed-out replies continue instead of failing
Claude Code stops losing whole turns to API timeouts and gets back drafts you clear by mistake.
- DeepSeek Harness Desktop — the open agent harness becomes a Mac and Windows app
DeepSeek's plugin-based agent harness now installs like a normal app on Mac and Windows.
- Claude Code 2.1.287 — Claude Mods let plugins change deeper behavior
Claude Code plugins can now reach deeper into how the agent behaves, starting with a side agent that watches your back.
- Codex CLI 0.160.0 — projectless sessions and Guardian review context
Codex CLI can now start without a project and gives its Guardian reviewer more context.
- Pi 1.0 — Earendil's minimal coding agent reaches its first stable release
Earendil calls Pi 1.0 a hardened, minimal, extensible agent harness that you can make your own.
- Magnitude — a local inference engine for agents, up to 2x faster than llama.cpp
An open inference engine that tunes itself to your machine so local agents run open models faster.
- Claude Code 2.1.286 — counted permission prompts and model-refusal retries
Claude Code numbers piled-up permission prompts, survives model refusals, and stops leaking secrets in logs.
- Pi 0.99 — the minimal coding agent adds MCP and a codemode sandbox
Pi's team used to say no to MCP. Version 0.99.0 adds it, with a JavaScript sandbox that composes tool calls.
- Claude Code 2.1.285 — a switch to turn off WebFetch and a --desktop handoff
Claude Code gets a kill switch for web fetching, a jump to the desktop app, and a long list of subagent and MCP fixes.
- ChatGPT Space and Pages — OpenAI's shared workspace for teams and agents
OpenAI puts documents, files and agents in one shared place inside ChatGPT, aimed straight at Microsoft Office.
- OpenAI Decisions API — GPT-6 Luna picks from your answers in 150 ms
OpenAI's answer to Jev: an endpoint that returns a choice, not text, about ten times faster than a normal Luna call.
- OpenAI Dots — always-on agents with their own cloud computer
OpenAI gives ChatGPT subscribers a background agent that keeps working on their goals when they are not looking.
- Manus 2.0 and Cue — agents with their own email, phone and wallet
Manus rebuilds its agent on a new harness and launches Cue, where each agent gets an email, a phone number and a wallet.
- Codex CLI 0.159.0 — steer the agent mid-response with instant interrupt
Codex CLI's new release lets you redirect the agent without waiting for it to finish.
- Claude Code 2.1.284 — Sonnet 5.5 becomes the default Sonnet with 1M context
The terminal agent switches to Anthropic's new Sonnet on the day it ships and starts showing costs in dollars.
- OpenRig — Claude Code and Codex agents run as one team from YAML
A control plane that turns a pile of coding-agent terminals into a named, persistent team.
- Codex CLI 0.158.0 — copy-on-select and MCP servers with OAuth secrets
Codex CLI's new release makes the fullscreen view easier to copy from and opens it to OAuth-protected MCP servers.
- Drawgent — your coding agent draws on a live Excalidraw canvas
A shared Excalidraw board where the coding agent you already use sketches and edits diagrams while you watch.
- Univer 1.0 — an open-source office SDK built as a harness for AI agents
One open-source runtime for office documents that both people and AI agents can edit.
- Claude Code 2.1.283 — admins can block models and audit old prompts
Admins get tighter control over which models run, and users can check their prompt files for habits written for older models.
- New Microsoft Copilot — one app for chat, code and always-on agents
Microsoft folds chat, a document-editing agent, an app builder and an always-on agent into one Copilot app.
- Ollaya — an Ollama-style runtime for local decision models
Ollaya does for typed decision models what Ollama does for LLMs: one binary to pull, run and serve them locally.
- Whiteboard — an open-source app for reviewing what coding agents built
A shared canvas where a coding agent draws diagrams of its changes, so you review the design before the line-by-line diff.
- Codex CLI 0.157.0 — GPT-6 Sol and Luna reach Amazon Bedrock users
Codex CLI's new release brings OpenAI's newest mid and low tiers to Bedrock and makes the fullscreen view the default.
- Claude Code 2.1.282 — project settings can no longer switch on telemetry export
A cloned repo can no longer quietly turn on telemetry export, and Claude Code now tells you which telemetry variables it ignored.
- Cursor Rollouts and Security Review — bots that watch a PR into production
Two new Cursor bots: one checks a change's health after it deploys, the other hunts exploitable bugs in every PR.
- Claude Code 2.1.281 — auto mode now asks before rm -rf "$(pwd)"
A risky rm pattern now always asks, attribution can be switched off, and resumed sessions keep their reasoning and cache.
- Ray-Ban Meta Audio and Gen 3 — Meta's new AI glasses from Connect 2026
Meta's Connect 2026 lineup adds camera-free audio glasses and a third-generation Ray-Ban Meta with a Meta AI button.
- LiteLLM v1.102.0 — guardrails finally run on streaming responses
Post-call guardrails stop being an end-of-response check and start rewriting text while it streams.
- Codex CLI 0.156.1 — GPT-6 Sol and Luna join the model picker
One day after OpenAI launched them, GPT-6 Sol and GPT-6 Luna are selectable inside the Codex CLI.
- Unreal Agent — an open harness that runs tool calls in the background
An MIT-licensed Go harness that keeps the model working while its tools finish.
- Codex CLI 0.156.0 — an optional fullscreen UI with transcript search
A fullscreen terminal UI, searchable transcripts and voice on by default land in Codex CLI 0.156.0.
- Claude Code 2.1.280 — Opus 5.5 arrives and Pro plans move off Sonnet
Claude Opus 5.5 becomes the default Opus model, and Pro and Team Standard plans now start on Opus instead of Sonnet.
- Pirate Face — open model weights turned into magnet links
Pirate Face mirrors open Hugging Face models as checksum-verified torrents, so the weights stay downloadable after the original is pulled.
- AX v0.3.0 — Google's agent orchestrator moves task state to Redis
Google's open agent orchestrator splits into three services and keeps task state in Redis instead of Kubernetes.
- OpenClaw 2026.9.5 — plugins hot-reload and GPT Live joins your calls
Plugins now install while the Gateway keeps running, and OpenClaw 2026.9.5 can sit in a meeting or a phone call with you.
- json-render 0.21.0 — Vercel Labs adds a TanStack Start renderer
Vercel Labs' generative UI framework now renders whole routed applications, not just single views.
- SGLang v0.5.20 — CUDA 12 wheels retired, radix cache covers every model
The September SGLang release trades CUDA 12 support for a smarter prefix cache, faster RL rollouts and nine more models.
- Claude Code 2.1.278 — auto mode's safety checks stop costing you tokens
Auto mode's pre-flight safety checks move to Anthropic's servers, and you are no longer billed for them.
- Google CC opens to families — one AI agent for up to six people
Google Labs' CC agent now belongs to a household rather than one person, with a shared morning brief, calendar and task list.
- Muse comes to Mac — Meta's agent acts inside your desktop apps
Muse now ships as a Mac app that works inside the native file, mail and calendar apps instead of only in a phone app or a browser tab.
- SemIf (formerly OpenJev) — typed decisions without generating JSON
An open take on the 'semantic if': read typed option probabilities from a small model instead of asking it to write JSON.
- Claude Code 2.1.277 — AGENTS.md works when there's no CLAUDE.md
One AGENTS.md can now brief Claude Code too, in any repo that has no CLAUDE.md.
- Claude Code 2.1.275 — claude.ai skills and plugins sync to the terminal
The skills and plugins you switch on at claude.ai now follow you into the terminal.
- Codex CLI 0.155.0 — voice conversations arrive in the terminal
Codex CLI can now be talked to: an experimental voice mode with live transcripts sits in the composer.
- SoL-Pi — NVIDIA's harness extension cuts coding-agent tokens by about half
Four efficiency tricks, picked by an automated research loop out of 152 candidates, packaged as a drop-in Pi extension.
- Astra for Law — OpenAI ties GPT-6 Astra to a 230M-URL legal index
OpenAI's GPT-6 Astra, wired to a searchable index of US law and 26 legal vendor plugins.
- Bend 2 — a language that makes an AI prove its code obeys your laws
Declare the rules your program must never break, and Bend 2 refuses to compile an edit that cannot prove it kept them.
- BrowserSkill 0.3.0 — Tencent's agent bridge gets canvas and remote gateways
Your agent borrows a window in the browser you are already logged into, then hands it back.
- Google Home MCP — any MCP agent can now run your smart home
Google's smart home platform now speaks MCP, so the agent you already use can read and control your Nest devices.
- Cowork folds into Claude — Anthropic adds Claude Docs and Claude Slides
Anthropic removed the line between Claude chat and Claude Cowork, and put document and slide editing inside the conversation.
- OpenAI Sponsored Agents — ChatGPT ads you can talk back to
ChatGPT ads become two-way: OpenAI is testing agents that answer questions on a business's behalf.
- Firefox Smart Window runs on Mistral Small 4 — and opens in France
Mozilla's browser assistant adds Mistral Small 4 as a model option and opens its beta to France.
- Claude Code 2.1.273 — a subshell could hide a dangerous rm in bypass mode
Claude Code 2.1.273 closes two permission-checker gaps and stops the context meter double-counting advisor-tool turns.
- Open Code Review v1.12.3 — secrets and .env files stay out of the review
Alibaba's open-source code reviewer now refuses to send secret paths and per-environment .env files to the model.
- Ollama 0.34.1 — MLX safetensors leave experimental, GGUF needs llama.cpp
Ollama 0.34.1 promotes MLX safetensors model creation out of experimental and hands GGUF conversion to llama.cpp tooling.
- dbt Charts — dashboards as YAML, so an agent can write them
One YAML file describes a whole interactive dashboard, so charts live in Git next to the models they read.
- Claude Code 2.1.271 — a sandboxed command only reaches its own hosts
Network permission moves from the whole session down to the single command that asked for it.
- LiteLLM v1.101.0 — smarter complexity routing and a semantic MCP search
The gateway gets a second-generation complexity router, a circuit breaker for its classifier, and semantic search over MCP tools.
- llama.cpp v0.4.1 — Maple 20B-A1B and Tencent Hy 4 now run locally
Three new model architectures land in the local inference engine, and three old loading flags are taken out.
- Siri AI ships in iOS 27 — Apple's rebuilt assistant goes live in beta
The rebuilt Siri is now on shipping software: onscreen awareness, personal context and systemwide app actions, in English first.
- Pion — Andon Labs opens a cloud platform where agents run a business
Andon Labs opens Pion, where persistent agents run a real company with a terminal, email, phone, banking and a browser.
- Lema AI Governance — third-party AI found, assessed and monitored
Lema AI Governance finds the AI a vendor added after you approved it, scores the exposure it creates, and watches it for drift.
- OpenClaw 2026.9.4 — plugins and skills install from the Control UI
OpenClaw 2026.9.4 gives plugins and skills one workspace in the Control UI, and lets a failed update roll itself back.
- Google ADK 2.9.0 — agents fail over to a backup model automatically
Agents keep running when a model errors, and ADK now speaks over the phone through LiveKit.
- Cursor Projects — a coordinator agent that delegates to thousands of subagents
A project-level agent in Cursor that plans, delegates to parallel subagents, and keeps working in the cloud after you close your laptop.
- Gemini for Windows — Google's desktop app opens over your work with Alt + Space
Google's Gemini assistant ships as a native Windows app, one Alt + Space away from whatever you are working on.
- Claude Code 2.1.269 — claude plugin eval scores a plugin against a baseline
A built-in eval runner for Claude Code plugins, with a no-plugin baseline that shows what the plugin actually contributes.
- DeepSeek Recipe — the official prompt encoder for V4 and V4.1
DeepSeek's own library for turning API requests into V4 and V4.1 prompts, and streamed output back into responses.
- Claude Managed Agents add 'auto' mode — the server checks every tool call
A third permission policy lets Anthropic's server decide, call by call, whether an agent's tool runs, stops, or waits for you.
- OpenAI Agents API — the Codex harness opens up to developers
The harness behind Codex, now a managed API: OpenAI runs the agent loop, your app brings the tools.
- Premiere's Generative Media Tool — five AI video models in the timeline
Premiere editors can now generate clips and sound effects on the timeline itself, choosing between Adobe's model and four rivals.
- Codex CLI 0.154.0 — GPT-6 Astra in the picker and git worktree sessions
Codex sessions can now run in their own git worktree, so parallel agents stop fighting over one checkout.
- Claude Code 2.1.267 — one setting caps effort on every provider
Claude Code 2.1.267 lets an admin cap effort level everywhere at once and spends most of its changelog repairing prompt-cache reuse.
- Desert Ant Labs — 18 small AI models that run offline inside your app
Eighteen small models that run fully offline inside iOS, Android and web apps, with no token bill.
- TeamAI CLI — Tencent ships one AI setup for a whole team
One git repo holds a team's agent skills, rules and MCP servers, and every member's coding agent pulls them automatically.
- Muse — Meta's personal AI agent runs in its own secure virtual machine
Meta's Muse takes an errand end to end — shopping, bookings, forms — from an ordinary chat thread.
- MCP Python SDK 2.2.0 — idle sessions now close after 30 minutes
The official Python SDK for MCP tightens session handling and OAuth checks, with the same fixes backported to the 1.x line.
- Humanizer v3.0.0 — the AI-writing cleanup skill drops 35 patterns to 25
An agent skill that strips the 25 habits which make text read as AI-written, without changing the facts.
- LiteLLM v1.100.0 — Vertex AI Interactions API and shared team budgets
The LLM gateway's newest release widens provider coverage and moves budget enforcement from single keys to shared groups.
- Ollama 0.34 — local models now run inside ChatGPT Desktop
Ollama 0.34 wires local open models into ChatGPT Desktop, set up from the Ollama app on macOS.
- OpenClaw 2026.9.2 — GPT-6 Astra support and swarms on by default
OpenClaw's September release picks up GPT-6 Astra on day one and switches concurrent sub-agent swarms on for everyone.
- llama.cpp v0.4.0 — Qwen3.8-Flash-Next support and lazy tensor loading
The local LLM runtime picks up two new model architectures and learns to read weights from disk on demand.
- LangChain 1.4.0 — a built-in MCP adapter for agent tools
LangChain agents can reach Model Context Protocol servers through a first-party adapter.
- SGLang v0.5.19 — beam search arrives, plus 786 merged pull requests
SGLang's September release adds beam search, nine more models, and an AMD attention kernel that fills idle compute units.
- Soup v0.74.0 — a dtype bug was doubling every fine-tune's memory
A missing dtype argument meant Soup loaded the frozen base model at twice its checkpoint precision, on every single fine-tune.
- Shunt — Spotify's Claude Code plugin cuts token use by 90%
A Claude Code plugin that hands bulk file reads to a cheap model, which Spotify measured at about 90% fewer tokens.
- Claude Code 2.1.261 — /skill-doctor shows which skills waste your context
Claude Code 2.1.261 adds a skill audit, much bigger inline output limits, and a stricter rm -rf safety check.
- Project HydraFusion — GitHub Copilot picks the model workflow for you
GitHub's router sends each coding task to the cheapest workflow that still clears the quality bar.
- Gemini Spark connects to Google Photos — it can edit, sort and share for you
Google Photos becomes a connected app for Gemini Spark, so one prompt can search, edit, album and share across a whole library.
- Qwen3.8-27B on Cerebras — 1,500 tokens per second at $0.99 per million
Cerebras added Qwen3.8-27B to its public endpoints, running Alibaba's 27B dense multimodal model at roughly 1,500 tokens per second.
- Claude Code 2.1.260 — a live diff panel and a permission-rule security fix
Claude Code 2.1.260 puts a live diff beside the conversation and repairs permission rules that quietly left folders writable.
- Claude Content Checker — see if a file carries Claude's signed credential
A free in-browser page that reads Claude's C2PA Content Credential out of a file's metadata.
- Codex CLI 0.153.0 — install plugins straight from remote marketplaces
Codex CLI 0.153.0 brings a plugin marketplace client, Vim undo and redo, and session reconnection that keeps your draft.
- Cursor self-hosted machines — cloud agents run inside your own network
Cursor Cloud Agents can now do their tool calls on hardware you control, while the model still runs in Cursor's cloud.
- Claude Code 2.1.259 — admins can push MCP servers to every user
The 2.1.259 release moves MCP configuration up to the administrator, and gives headless runs a way to refuse prompts instead of hanging.
- Slotstream v0.2.0 — a 104GB model on a 48GB Mac, plus speculative decoding
Slotstream streams a 125B model's experts off SSD so a Mac with 48 GB of RAM can run a 104 GB checkpoint.
- Google Pics — an AI image tool in Workspace where you start from a prompt
Google Pics turns a written description into a poster or graphic, then lets you edit one object at a time.
- ChatGPT connects to Epic — clinicians can pull chart context into the chat
Hospitals can wire their Epic records into ChatGPT, and clinicians get a plugin for nine official public health data sources.
- Agentic video in Gemini — the model loads only the clips it needs
Gemini can now walk a video on its own, pulling only the segments a question needs instead of every frame at a fixed rate.
- Claude Code 2.1.257 — Claude Fable 5.1 becomes the default Fable model
Claude Code 2.1.257 switches its default Fable model to Claude Fable 5.1 and stops auto mode waving through container-escape moves.
- ChatGPT Mil and Grok for Government — the Pentagon's AI portal adds two models
The Pentagon's GenAI.mil portal now runs ChatGPT Mil and Grok for Government next to Google's Gemini.
- Codex CLI 0.152.0 — the planning tool is now off by default
Codex CLI 0.152.0 turns the planning tool off by default and tightens how MCP tools return output.
- OpenClaw 2.0 — the largest release yet for the open-source AI assistant
OpenClaw 2.0 rebuilds installation, memory, skills, the browser and team access in one release built from over 16,000 merged pull requests.
- Claude for Teachers reaches districts — a free Enterprise plan for U.S. K-12
Schools can now hand out Claude for Teachers centrally instead of asking each teacher to sign up alone.
+ 358 more in the sitemap.