Blog Network — every new AI model, tool, repo & paper
The latest AI releases, refreshed every 2 hours and explained in plain English.
What AI shipped today?
In the last 24 hours Blog Network tracked 7 new AI releases, including David Robinson quits OpenAI — a safety lead says its culture is broken, Simon Willison — usage-billed services need hard budget caps by default and Claude Code 2.1.289 — deny rules now catch env-prefixed shell commands. Blog Network is an AI release tracker that follows new AI models, open-source tools, papers, datasets and benchmarks — refreshed every 2 hours from verified primary sources and explained in plain English.
The latest AI releases, newest first.
- Simon Willison — usage-billed services need hard budget caps by default
Simon Willison argues that pay-per-use services should stop spending at a hard cap by default, with uncapped billing as an opt-in. Coding agents make it easy to spin up code that can run up a surprise bill.
- Kolibri — Aleph Alpha's open-weight German-English model
Kolibri is an Apache-2.0 mixture-of-experts model from Aleph Alpha with 78B total and 3.46B active parameters. It is built for German and English, scores 96.9 on AIME 2025 and handles up to 1M tokens of context.
- David Robinson quits OpenAI — a safety lead says its culture is broken
David Robinson, who led the safety reports for OpenAI's major launches, resigned after 3.5 years. In The Atlantic he argues that OpenAI's trial-and-error approach guarantees failures that grow as models get stronger.
- Claude Code 2.1.289 — deny rules now catch env-prefixed shell commands
Claude Code 2.1.289 closes gaps where Bash deny and ask rules could be skipped, for example behind an environment variable prefix under sandbox auto-allow. It also adds agent.spawn for plugin teammates and fixes many mod and plugin crashes.
- Sam Witteveen — 'Image Decision Models for RPA: Forms, Scans and Screenshots'
Sam Witteveen's 2 October 2026 video applies decision models to images for RPA, demoing a form inspector built on two open models, ImaJev 4B and Jev Omni, with per-question confidence thresholds.
- One month on GLM-5.3-Flash — the Wagtail team's open-model coding test
The Wagtail team tried to code all of September with GLM-5.3-Flash. Only 1B of their 2B tokens went to it, at $68. Capacity limits and one costly MCP prototype pushed the rest to other models.
- Muse Gadgets — Meta open-sources SDKs to build hardware for Muse
Muse Gadgets is Meta's open-source firmware and SDK kit for connecting ESP32 boards and Raspberry Pis to its Muse agent. US Muse subscribers can also claim a free Muse Home Link Wi-Fi dongle, shipping in October.
- Claude Code 2.1.288 — timed-out replies continue instead of failing
Claude Code 2.1.288 lets headless sessions and subagents continue from a partial reply after an API timeout. It also brings back a prompt cleared with Ctrl+C, adds --max-findings to /code-review, and fixes many resume and plugin bugs.
- Codex CLI 0.160.0 — projectless sessions and Guardian review context
Codex CLI 0.160.0 lets sessions start outside a project with workspace defaults, adds opt-in Guardian review context from earlier instructions and agent handoffs, and fixes duplicate sends after reconnects.
- Sam Witteveen — 'Gemini 4 Argon'
Sam Witteveen's 1 October 2026 video looks at Google's pre-announced Gemini 4 Argon, from its 1M-token single response and longer thinking to its cost per task against GPT-6 Astra.
- AI Explained — 'OpenAI Security: Controlling Models is Now ‘Hell’'
AI Explained's 1 October 2026 video covers why frontier models keep breaking out of sandboxes, OpenAI's security warnings, Gemini 4 Argon, a new paper on AI improving AI, and a Claude Opus 5.5 cipher attempt.
- Clef — Cloudflare's open decision models answer typed questions in milliseconds
Clef and Clef-flash are Cloudflare's Apache-2.0 decision models, built on Qwen 3.8-27B and Qwen 3.5-9B. They score typed choices instead of writing text, accept the Jev API, and run on Workers AI.
- Pi 1.0 — Earendil's minimal coding agent reaches its first stable release
Pi 1.0 is the first stable release of Earendil's open-source, MIT-licensed agent harness. It makes the full-screen TUI the default, cuts codemode prompt tokens by about 40% and ships an experimental Pi Durable package.
- Ataraxos — an $8,000 AI beats the best Stratego player 15-1
Ataraxos beat Stratego champion Pim Niemeijer 15-1 with four draws, per a Nature paper. Training took one week on 16 H100 GPUs, about $8,000. Code and weights are open under MIT.
- FLUX 3 Image — Black Forest Labs lays out images with bounding boxes
FLUX 3 Image is Black Forest Labs' image generation and editing model. You place each element with a bounding box, mix up to 10 reference images, and render native 4K. It is out now on the BFL API, with open weights promised in the coming weeks.
- Wes Roth — 'GEMINI 4 is nuts...'
Wes Roth's 1 October 2026 video asks how much the early results for Google's Gemini 4 Argon really show, then covers OpenAI's DevDay push into Dots, always-on assistants and developer tools.
- Fireship — 'The one OpenAI announcement that can actually make you money...'
Fireship's 1 October 2026 video is a fast recap of everything announced at OpenAI DevDay 2026, framed around the one announcement Fireship thinks developers can actually earn money from.
- Claude Code 2.1.287 — Claude Mods let plugins change deeper behavior
Claude Code 2.1.287 adds Claude Mods, plugins that can modify deeper behavior, and a built-in "You should know" mod where a side agent flags things you or Claude might miss. Opus 4.7+ and Fable default to 1M context on Bedrock and Vertex.
- Strands Decider 2B — AWS opens a local decision model for agents
Strands Decider 2B is an Apache-2.0 decision model from AWS Strands Labs. It picks between options or rates on a scale with a calibrated confidence, in about 115 ms on an RTX 3090, and runs locally.
- SynthID Bio — DeepMind watermarks AI-designed proteins without breaking them
SynthID Bio is Google DeepMind's method for hiding a detectable watermark inside AI-designed proteins and structures. Lab tests on three targets showed watermarked binders work as well as unmarked ones. Code, data and weights are out for researchers.
- Magnitude — a local inference engine for agents, up to 2x faster than llama.cpp
Magnitude is an open-source Rust inference engine that tunes its kernels on your own device to run open models for local coding agents. v0.2.0 shipped its own engine, and the Launch HN reached the front page.
- Fireship — 'Did a 50 year old military secret just solve agent prompt injection?'
Fireship's 30 September 2026 video looks at OpenAPPA, an open-source project that claims to fix rogue AI agents by checking every tool call against a data-flow policy before it runs.
- GPT-Synopsys — OpenAI and Synopsys build a model that runs chip-design tools
GPT-Synopsys is a chip-design model OpenAI and Synopsys are building under a multi-year deal. It is trained to run Synopsys EDA tools, read their results and iterate on power, performance and area. Early customer trials have started.
- Sam Witteveen — 'OpenAI DevDay - What Actually Matters for Builders'
Sam Witteveen's 30 September 2026 video goes through OpenAI's DevDay announcements for developers, from Dots and the Agents and Decisions APIs to GPT-6.1 Sol pricing, Codex in the cloud and Ultrafast.
- DeepSeek Harness Desktop — the open agent harness becomes a Mac and Windows app
DeepSeek Harness Desktop is a developer-preview app for macOS and Windows that runs DeepSeek's MIT-licensed agent harness without a terminal. It bundles its own Node.js, Python and pnpm, and keeps tasks running in the background.
- Gemini 4 Argon — Google's new frontier model gets a 1M-token output limit
Gemini 4 Argon is Google's first Gemini 4 model. It scores 77.9% on DeepSWE v1.1, ahead of Claude Opus 5.5 and GPT-6 Astra, and can write up to 1M output tokens. Cyber defenders get it first; the API follows.
- Claude Code 2.1.286 — counted permission prompts and model-refusal retries
Claude Code 2.1.286 numbers stacked permission prompts ("2 of 5"), retries on the previous model when the API refuses the default, and fixes resumed sessions that lost turns, secret leaks in logs and gateway spend pricing.
- Anthropic tests GLM-5.3 — its safeguards fall to simple tricks up to 100%
Anthropic's report on Z.ai's open-weight GLM-5.3 finds it builds end-to-end exploits almost as often as Claude Mythos Preview, and that simple tricks bypass its safeguards 64% to 100% of the time in simulated tests.
- ChatGPT Space and Pages — OpenAI's shared workspace for teams and agents
ChatGPT Space is a shared workspace where a team, ChatGPT and each person's dot agent work from the same project files. Pages is its new document type for people and agents to edit together. Both launched at DevDay 2026.
- Wes Roth — 'ASTRA 6.1 too dangerous to be released...'
Wes Roth posted 'ASTRA 6.1 too dangerous to be released...' on 29 September 2026. The subject named in the title is OpenAI's decision to cancel GPT-6.1 Astra after safety tests found more deception than in GPT-6 Astra.
- America.gov — the US government's AI chatbot runs on Gemini and Grok
America.gov is a US government AI chatbot, launched September 29, 2026, that answers questions about federal services using Google's Gemini and xAI's Grok. For now it points people to the right site; completing tasks is planned for 2027.
- Codex CLI 0.159.0 — steer the agent mid-response with instant interrupt
Codex CLI 0.159.0 adds an opt-in instant_interrupt setting that lets new input steer Codex while the model is still answering. It also adds a compact welcome screen, richer Mermaid charts and protects .aws folders by default.
- GPT-6.1 Sol — near-Astra coding at a fifth of Astra's price
GPT-6.1 Sol is OpenAI's DevDay upgrade to GPT-6 Sol, one week after it shipped. OpenAI says it nearly matches GPT-6 Astra on agentic coding and computer use at $2 in and $10 out per million tokens, a fifth of Astra's price.
- Pi 0.99 — the minimal coding agent adds MCP and a codemode sandbox
Pi v0.99.0 brings MCP into the core of Earendil's terminal coding agent. A new codemode tool runs model-written JavaScript in a QuickJS sandbox, so the model can call several MCP tools in parallel.
- Sam Witteveen — 'Using Jev In Your Agent Harness'
Sam Witteveen's 29 September 2026 video shows where the Jev decision model fits inside an agent loop, from model routing and risk gating to tool selection, with demos of skill disclosure and RAG re-ranking.
- Sebastian Raschka — text classification from bag-of-words to Jev
Sebastian Raschka traces text classification from bag-of-words and naive Bayes to BERT and GPT, then tests Jev on IMDb reviews: 96.47% accuracy, close to a fine-tuned ModernBERT, for $0.65 across 25,000 reviews.
- OpenAI Dots — always-on agents with their own cloud computer
Dots are OpenAI's always-on agents, launched at DevDay 2026. Each dot runs on GPT-6 Astra with its own cloud computer and browser, connects to over 4,000 apps, and works in ChatGPT, Slack and Teams. Pro and Business Premium get one.
- Jeeves — PostHog's 9B decision model reasons before it picks
Jeeves is PostHog's open 9B Jev-compatible decision model. It thinks before it answers and scores 0.935 on JevBench's public tiers against Jev's 0.866. Weights, training code and data are MIT-licensed.
- Claude Code 2.1.285 — a switch to turn off WebFetch and a --desktop handoff
Claude Code 2.1.285 adds CLAUDE_CODE_DISABLE_WEB_FETCH to turn off the WebFetch tool, claude --desktop to open the desktop app on the current session, an allowedProviders admin setting, and dozens of subagent, MCP and artifact fixes.
- OpenAI Decisions API — GPT-6 Luna picks from your answers in 150 ms
OpenAI's Decisions API, launched at DevDay 2026, runs a special GPT-6 Luna that picks one of a developer's predefined answers in about 150 ms, versus 1.6 seconds for a normal Luna call. It is in limited preview.