New AI Tools — Products, CLIs & IDE Features | Blog Network
The newest AI tools, products, CLIs and IDE features — what shipped, who it's for, and how to try it tonight, in plain English.
344 releases tracked
- Claude Code 2.1.243 — the install drops from 340 MB to 75 MB
A lighter Claude Code: 75 MB to install, 40-70 MB less memory per session, plus four new settings for organizations.
- Tempus ECG-PH — FDA clears AI that spots pulmonary hypertension in a routine ECG
An FDA-cleared AI that reads a routine ECG and flags patients who may have pulmonary hypertension.
- NVIDIA Groq 3 LPX — the agent inference chip enters full production
NVIDIA's Groq 3 LPX accelerator is in full production, built to decode tokens fast enough to keep AI agents responsive.
- MoneyPrinterTurbo v1.3.5 — Claude joins the one-click short-video maker
MoneyPrinterTurbo builds a finished vertical short from a keyword, and v1.3.5 adds Claude, two more voice engines and API-key auth.
- LiteLLM v1.98.0 — reserved capacity gets flat-cost billing, not per-token
LiteLLM v1.98.0 teaches the open-source AI gateway to bill reserved capacity and to test routing changes before adopting them.
- Claude Code 2.1.239 — a proxy bug that doubled Bedrock API calls is fixed
Anthropic's 2.1.239 build closes a Bedrock proxy bug that silently doubled the API calls you were billed for.
- SGLang v0.5.18 — cold starts get 2.38x faster, seven model families land
SGLang v0.5.18 overlaps weight loading with CUDA graph capture, cutting a large-model cold start from 85 seconds to 36.
- Bot Preference Sync — Cloudflare writes your robots.txt to match your bot rules
Cloudflare builds your robots.txt out of the AI bot policy you already set in its dashboard.
- LLM 0.33 — Simon Willison's CLI moves to the OpenAI Python 3.x library
LLM 0.33 rebases Simon Willison's model CLI on the OpenAI Python 3.x library and adds per-call embedding keys.
- Claude browser use tool — Anthropic's agent APIs are now generally available
Anthropic's agent building blocks leave preview, joined by a browser tool that targets page elements instead of pixels.
- OpenAI regional processing — pick an inference region per API request
One API key, ten regional domains — choose where each OpenAI request is processed.
- NoBuzz — a Claude Code skill that rewrites Claude's replies in plain English
NoBuzz hands Claude's answer to a second model and prints the plain-English rewrite instead.
- Codex CLI 0.149.0 — a dashboard for every running agent task
Codex CLI 0.149.0 puts every agent task behind one searchable dashboard you can drive with shortcuts.
- Claude Code 2.1.238 — plugin marketplaces can mint their own auth headers
Private plugin marketplaces can now hand Claude Code a fresh token for every catalog fetch.
- Waymo's custom AI chip — a 5nm ASIC that runs 1,000+ TOPS in the car
Waymo now designs the silicon that turns its robotaxi's sensor data into driving decisions.
- Slack Code — AI coding agents get a channel the whole team can watch
Slack Code gives a coding agent its own project channel, so planning, diffs and sign-off all happen in front of the team.
- Antigravity IDE Extensions — Google's agent moves into VS Code and JetBrains
Google's agent-first coding platform now installs as a normal extension in the editor you already use.
- Apple Messages plugin — ChatGPT can read and send texts on a Mac
ChatGPT can now open your Mac's Messages app, search old threads and send a reply you approve.
- Huzzah — an editor where you write pseudocode and an LLM fills in the code
Write pseudocode instead of a paragraph of English, and Huzzah syncs it into working code.
- Vomit — a local model rewrites Claude Code's replies before you read them
Vomit pipes Claude Code's output through a small local model so the prose comes back plain.
- Ramp Router — an LLM gateway that cut customer inference bills 40% on average
Ramp Router takes one API call and forwards it to the model that meets your quality bar at the lowest price.
- NeuroQuant PET — FDA clears automated amyloid scoring for dementia
Cortechs.ai's NeuroQuant PET can now be used on real patients to score amyloid brain scans automatically.
- Mistral Agentic Search — models search, open and grep their way through docs
Mistral's retrieval layer lets an agent search, open and grep a document set until it can actually answer.
- Meta AI for Mac — a native desktop app with screen sharing and dictation
Meta's assistant gets a real Mac app that can read your screen and type for you anywhere.
- Binance Agent OS — an MCP server that lets AI agents place real trades
Binance opens its trading, wallet and payment stack to AI agents through a single MCP server.
- ChatGPT Ads reach Europe — 31 markets get labeled ads on Free and Go
OpenAI takes ChatGPT Ads into 31 European markets, with GDPR consent choices and no ads on paid plans.
- ai-memory v1.29.0 — long-term memory that follows agents across CLIs
A local, git-versioned wiki of what your coding agent already figured out, readable by whichever agent you open next.
- fx — Vercel Labs open-sources a 6 MB coding agent written in Zig
A tiny native coding agent from Vercel Labs, built to be embedded anywhere a shell can run.
- Gemini study tools — 3D simulations and a free year of Google AI for students
Google gives verified college students a free year of its paid AI plans and adds study tools across Gemini, Search and Lens.
- Cursor Subscriptions — cloud agents watch a PR and drive it to done
Cursor's cloud agents can now subscribe to a PR, a Slack thread or a schedule and keep going on their own.
- oMLX 0.6.2 — the Mac LLM server now tunes its own ANE/GPU split
Instead of shipping one ANE/GPU ratio for every Mac, oMLX now measures the best split on the machine in front of you.
- Palomar — a registry that machine-checks Lean proofs, human or AI
A searchable registry of Lean formalizations whose proofs are replayed through two independent kernels before they are listed.
- Claude Playground — Anthropic retires Workbench for an API-exact console
Playground replaces Workbench in the Claude Console and shows you the exact API request your prompt turns into.
- Warp Factories — cloud agent pipelines for the whole dev cycle
Warp Factories turns a backlog ticket into a reviewed pull request using fleets of cloud coding agents.
- Claude Code keeps 50% higher weekly limits — extended through August 31
Claude Code's 50% weekly limit boost, first added in May, now runs through August 31, 2026.
- ChatGPT for Teens — OpenAI's age-gated mode for 13-to-17-year-olds
OpenAI's teen mode locks down sensitive topics, steers homework into Study Mode, and hands parents a Quiet Hours switch.
- Cursor Origin — a Git forge for agents opens in early beta
Cursor now hosts your code as well as writing it, with pull requests and GitHub sync in the same tab.
- Claude Code 2.1.233 — GitLab merge requests, marketplaces and token redaction
GitLab teams get first-class Claude Code support: merge request worktrees, plugin marketplaces and token redaction.
- Gemini watermarks become optional — Google adds an off switch for AI media
Google makes Gemini's visible AI watermark a user choice, while invisible SynthID and C2PA provenance stay in every file.
- Computer History — ChatGPT builds a memory from your Mac activity
ChatGPT for Mac can now build memories from what you do in your apps, using interaction events instead of screenshots.
- Credentio — Google open-sources the C++ library behind its content credentials
Google's C++ library checks C2PA Content Credentials on the device, with nothing sent to a server.
- HEIR — Google's compiler runs AI models on encrypted data
Google's HEIR compiler takes a normal trained model and rebuilds it to run on data the server can never read.
- Suno Studio 2.0 — browser music workstation adds MIDI and a chat bar
Suno's browser music workstation now takes MIDI, and a chat bar builds instruments, effects and synth presets on request.
- Cursor Builds — cloud agents fork a warm dev environment instead of setup
Cursor keeps warm copies of your development environment ready, so a cloud agent skips setup and starts working almost immediately.
- Ultrafast mode — GPT-5.6 Sol at 750 tokens per second on Cerebras
OpenAI's new API tier runs GPT-5.6 Sol on Cerebras hardware at up to 750 output tokens per second.
- DeepSeek Harness — open-source agent framework built entirely from plugins
An MIT-licensed agent framework from DeepSeek where models, tools, sandboxes and even the interface are swappable plugins.
- Unsloth Desktop — run and train local AI models without writing code
A free desktop app that runs, fine-tunes and serves local AI models on Windows, macOS and Linux.
- Compliance API pulls local sessions — admins can read Claude Code transcripts
Claude Enterprise admins can now pull transcripts of Cowork and Claude Code sessions that ran on a user's own laptop.
- Delta — Zed's multiplayer workspace for coding with agents
A multiplayer environment for coding with agents and for reviewing what those agents build.
- Orca 1.4.180 — stacked pull requests come to the parallel-agent desktop
Run Claude Code, Codex and other CLI agents side by side in isolated git worktrees from one desktop app.
- Grok Bot — xAI's always-on agents get their own cloud computer
xAI's agents now hold their own cloud machine, log into your tools, and finish the job while you are away.
- ChatGPT Desktop for Linux — OpenAI ships a preview with Codex built in
OpenAI's desktop app now runs on Linux, bringing ChatGPT, ChatGPT Work and the Codex coding workspace to Ubuntu, Debian and Fedora.
- Mojo 1.0 — Modular's AI systems language reaches its first stable release
Modular's Python-like language for GPU and AI systems code reaches 1.0 after three years of breaking changes.
- Metal Capability Shim — llama.cpp runs up to 16x faster inside macOS VMs
A process-scoped Metal shim unlocks the fast GPU kernels for llama.cpp inside macOS VMs on Apple Silicon.
- Mistral Regional Endpoints — pin inference to Europe or the US
Mistral API calls can now be pinned to European or US data centres, and outside open models run on the same platform.
- NeMo Switchyard — NVIDIA's open router picks a model per agent step
An Apache-2.0 Rust proxy from NVIDIA that reshuffles which model handles each step of an agent run.
- ChatGPT books restaurant tables — OpenTable, Resy and Yelp in the chat
Restaurant reservations and waitlists now happen inside a ChatGPT conversation instead of on a separate booking site.
- ChatGPT Business Premium seats — $125 a month for 5x usage, no 5-hour cap
OpenAI adds a $125-a-month Premium seat to ChatGPT Business, with five times the usage of a Standard seat.
- OpenChamber 1.18.2 — agent workspace adds scheduled tasks and a live panel
OpenChamber 1.18.2 puts an agent's goal, subagents and context usage in one live panel, and lets projects schedule recurring agent tasks.
- SGLang v0.5.17 — day-0 serving for Kimi K3 and MiniMax H3
SGLang v0.5.17 serves two frontier open models on release day and starts replacing its Python front-end with Rust.
- Claude Managed Agents get spend caps — a session pauses at its dollar budget
Claude Managed Agents sessions can now carry a hard dollar cap that pauses the agent instead of letting it keep spending.
- Claude Code cross-session messaging — one session can message another
One Claude Code session can now send a short written message to another, instead of you re-explaining the same thing in each terminal.
- Claude Code auto mode becomes the default — a classifier replaces most prompts
Claude Code stops asking before most tool calls and routes them through a safety classifier instead.
- Claude Code self-hosted environments — Anthropic runs sessions on your own compute
Claude Code cloud sessions can now run on runners you provision, keeping code and secrets on your infrastructure.
- Hark Handoff — browser-use agent predicts the next click, not the next token
Hark's first agent runs in its own virtual computer and is trained to predict the next click instead of the next token.
- Kitesurf — Cloudflare's agent-first browser runs in V8 isolates on Workers
A browser built for AI agents, not humans — Kitesurf runs in Workers and slashes CPU and memory versus Chromium.
- Google Ask Maps — Gemini agent orders food, books hotels, and buys event tickets
Google Maps' Ask Maps assistant can now order food, book hotels and buy event tickets in one sentence.
- Astro triagebot-action — AI issue triage drops Astro's open-issue count 85%
Astro's team shipped a GitHub Action that pushes AI agents to reproduce, debug, and verify issues before a human sees them.
- Zed 1.14 Sandboxing — OS enforces what agents can touch
Zed 1.14 makes the OS the wall around your coding agent, not the agent's promise to behave.
- Prime Agent — self-improving coding harness beats humans on ARC-AGI-3
An MIT-licensed coding harness that rewrites itself mid-task and edges past human experts on ARC-AGI-3.
- Castform — RL post-trains a 4B open model to match GPT-5.6 Sol on search at 1/100 the cost
A managed RL post-training service that specialises 4B open models to match frontier accuracy on narrow tasks for a fraction of the cost.
- Cloudflare OS — open agent workspace for building internal apps on your company data
Cloudflare's open-source answer to internal chat-plus-app platforms, built on Workers and released under Apache-2.0.
- Cloudflare Computer — agent runtime routes work between isolates and containers
Cloudflare's new agent runtime pairs a Durable-Object virtual filesystem with a fast isolate and a fallback Linux container.
- LLM 0.32 — Simon Willison's CLI ships reasoning traces and server-side tools
The de-facto Python CLI for LLMs turns its 0.32 alpha into a stable release, with reasoning traces, provider tools, and a Git-style log store.
- Warp Agent CLI — the Warp coding agent goes standalone in any terminal
Warp's multi-model coding agent ships as a standalone CLI you can drop into any terminal, with multiplexing, remote handoff, and BYO keys.
- Cloudflare Workers AI — FP8 doubles Kimi K2.6 context, INT4 cuts GLM 5.2 40%
FP8 KV cache and INT4 weights land on Workers AI, doubling context and shrinking checkpoints for Kimi K2.6 and GLM 5.2.
- Artifacts Hub — Interconnects launches open-model tracking dashboards
Interconnects ships two free tools that track open-model releases and adoption across geographies and organizations.
- MiniMax H3 Day-0 Support in ComfyUI — 2K video generation on an RTX 3060
Day-0 ComfyUI workflows plus a pruned + quantized H3 that fits on an RTX 3060.
- Turbo Fieldfare — Gemma 4 26B runs in about 2 GB of RAM on any M-series Mac
A 26-billion-parameter Gemma 4 model on an 8 GB MacBook Air — without swap.
- OpenAI Codex Security — Apache-2.0 CLI scans repos for vulnerabilities
OpenAI's new Apache-2.0 tool for finding, tracking, and gating security bugs in a codebase, from the command line or inside CI.
- Grafana AI Week — six agentic ops tools land in Grafana Cloud
Grafana's Assistant becomes an agent stack — investigations, MCP, and a CLI drop to GA on the same day.
- NVIDIA Agent Toolkit — PhysicsNeMo and CUDA-X plug into agents for chip and physics work
NVIDIA's Agent Toolkit gains PhysicsNeMo and CUDA-X libraries as agent-ready skills, aimed at automating chip design and physics workflows.
- Cohere North Automations — plain-English AI workflows across enterprise systems
Cohere adds a workflow orchestrator to its enterprise North platform, so teams can chain AI agents across their own systems from a plain-English brief.
- NVIDIA ModelExpress — Rust sidecar makes vLLM cold-loads 48× faster
NVIDIA's open-source Rust sidecar loads model weights into GPUs up to 48× faster than a cold Hugging Face pull.
- Meta AI can now act — Muse Spark 1.1 plans, uses your calendar, finishes tasks
Meta AI stops answering and starts doing — plans, tools, and follow-through, live in the app you already use.
- OpenRouter Classifiers — auto-tag every AI generation for cost tracking
OpenRouter runs a small model over every API generation to tag department, task type, and cost center — visibility without latency.
- Alibaba open-code-review — line-level LLM code review at 1/9 the tokens
Hybrid pipeline + LLM code reviewer, battle-tested at Alibaba, now Apache-2.0 with 15.7k GitHub stars.
- Grok in Google Workspace — free xAI add-on lands inside Docs, Sheets, and Slides
xAI turned Grok on inside every Google Workspace doc, sheet, and slide — free, one install, cited cells and all.
- Grok Build Workflows — xAI's coding CLI now fans a task across up to 1,024 parallel agents
Grok Build learned to fan a job across up to 1,024 parallel agents, verify with independent skeptics, and post one report — from a single slash command.
- Claude Code 2.1.219 — Opus 5 becomes default, subagents nest to depth 3
Anthropic's terminal coding agent adopts Opus 5 as its default and lets subagents spin up their own subagents three layers deep.
- AMD Helios — 72-GPU rack-scale AI system aimed straight at Nvidia's NVL72
AMD's first rack-scale answer to Nvidia's NVL72: 72 MI455X GPUs, 31 TB of HBM4, and hyperscaler commitments already signed.
- ChatGPT Voice lands on desktop — OpenAI puts GPT-Live in the Mac and Windows apps
ChatGPT Voice, powered by OpenAI's new GPT-Live model, is now inside the Mac and Windows apps.
- OpenAI Hard Spend Limits — API caps that stop requests, not just alert
OpenAI is finally letting every API account set a monthly bill ceiling that actually stops requests instead of just paging you.
- Claude voice mode adds Opus and Sonnet — Anthropic ends its Haiku-only voice era
Anthropic's voice mode can finally reach beyond Haiku to Opus and Sonnet, so voice conversations get the same models people already trust for text.
- ChatGPT Health opens to all US users — Apple Health and Epic records land in the chat
ChatGPT can now read your Apple Health data and Epic medical records — and use them to answer any question, not only inside a dedicated health hub.
- Substack ships AI-writing detection — Pangram scores every post and comment
Substack readers can now check whether a post, comment, or Note was written by a human, by AI, or somewhere in between — powered by Pangram.
- Echo — Tracer's router blends open-weight LLMs at a third of Fable 5's cost
One endpoint, many open-weight models, roughly a third of Claude Fable 5's inference cost on the tasks Tracer measured.
- Runway Media Router — one API picks the best image, video, or audio model per request
Runway ships the first preference-optimized router for generative media so developers stop hand-picking between video, image, and audio models.
- Cursor Router — automatic model picker chooses frontier or Grok 4.5 per request
Cursor's new Router picks a frontier model or the cheap Grok 4.5 for every prompt so users stop guessing which model to run.
- OpenAI Presence — enterprise platform to build, guard, and improve AI agents
OpenAI Presence is a managed platform for putting voice and chat agents in front of real customers, with guardrails, evaluations, and a Codex-powered improvement loop.
- GigaToken v0.9.0 — a Rust tokenizer that runs ~1000× faster than HuggingFace
GigaToken hits ~24.5 GB/s on a 144-core EPYC and drops straight into HuggingFace and tiktoken code paths.
- Anthropic Economic Index Connector — query AI labor data inside Claude
Anthropic's new Claude connector turns the Economic Index into a chat interface — pick a question, get a sourced answer.
- Buzz — Block's open workspace where humans and AI agents share channels
Block's open-source workspace where humans and AI agents share the same channels, code review, and workflows on a Nostr relay.
- Nativ v0.0.1 — local AI Mac app that discovers and monitors MLX models
A free, open-source Mac app that turns any MLX model on your disk into a local chat + OpenAI-style API.
- Unsloth v0.1.50-beta — AMD GPU support lands across Windows, WSL, and Linux
Unsloth's LLM training and inference stack now runs natively on AMD GPUs across Windows, WSL, and Linux.
- Grok Automations — scheduled and email-triggered jobs land in Grok apps
Grok can now run a saved prompt on a timer or fire it off the moment a matching email lands in your inbox.
- Roblox Build — mobile AI turns a text prompt into a playable game
A text prompt on your phone now compiles into a playable Roblox game, alpha on July 28 in New Zealand.
- Google Vids Personal Avatars — Gemini Omni videos from a selfie and voice clip
Google Vids becomes the first mainstream Workspace app to auto-cast the user as the on-screen presenter in an AI video.
- Google AI Mode adds Connected Apps — Instacart, Canva, and YouTube Music
AI Mode graduates from answering questions to doing things — Instacart carts, Canva designs, and YouTube Music playlists all inside Search.
- Claude Code 2.1.212 — /fork forks to a background session, agents get budgets
Claude Code's latest release makes /fork a background-session brancher and gives long-running tool calls and subagents hard budgets.
- LM Studio Bionic — new agent app built for open-source models
LM Studio ships a standalone agent app aimed at open-weight models, with local voice, code, and document workflows.
- Gemini Notebook — Google renames NotebookLM and gives every notebook a cloud computer
NotebookLM's rebrand ships a sandboxed cloud computer so notebooks can write and run code against your uploaded sources.
- Apple opens new Siri AI to the public — iOS 27 public beta arrives
Apple's rebuilt Siri, first unveiled at WWDC in June, finally reaches non-developers through the iOS 27 public beta.
- Talk to Spotify — conversational AI beta lands for Premium users
Spotify Premium users can now hold a real chat with the app to pick songs, learn about albums, and revisit their listening history.
- Codex Micro — OpenAI's first hardware is a $230 macro pad for Codex
OpenAI's first shipping hardware is a keypad for its Codex coding agent.
- Juggler — a GUI coding agent from the creator of JUCE
A visual, tree-based coding agent from a 30-year veteran of desktop developer tools.
- Claude for Teachers — free Claude built for K-12 educators
A free Claude tier for verified US K-12 teachers, wired to state academic standards and nine classroom platforms.
- Destructive Command Guard — Rust hook stops AI agents from rm -rf
A Rust hook that blocks AI coding agents from running commands like rm -rf ./src or git reset --hard.
- Mesh LLM — distributed inference on iroh's peer-to-peer network
Pool GPUs across your machines and serve any of 40+ models through one OpenAI-compatible endpoint at localhost:9337.
- OpenClaw 2026.7.1-beta.5 — conversational setup and GPT-5.6 support land
OpenClaw's beta.5 turns first-run setup into a live agent conversation and wires GPT-5.6 into the whole stack.
+ 224 more in the sitemap.