Blog Network — every new AI model, tool, repo & paper
The latest AI releases, refreshed every 2 hours and explained in plain English.
What AI shipped today?
In the last 24 hours Blog Network tracked 13 new AI releases, including Codex CLI 0.162.0 — managed Git worktrees and pinned tasks, Fired OpenAI safety researchers publish an open letter — they deny a leak and bigarrow — AI agents point at what to click with big arrows on your Mac. Blog Network is an AI release tracker that follows new AI models, open-source tools, papers, datasets and benchmarks — refreshed every 2 hours from verified primary sources and explained in plain English.
The latest AI releases, newest first.
- Wes Roth — 'OpenAI's "Alien Math" Is Freaking People Out'
Wes Roth's 9 October 2026 video covers the backlash to OpenAI's math release, a mathematicians' boycott call, crypto security worries from Vitalik Buterin and Justin Drake, and Anthropic's new usage policy.
- Anthropic Usage Policy 2026 — new rules for surveillance, robots and abuse
Anthropic updated its Usage Policy for Claude, effective 12 November 2026. It bans tracking people without consent, sets safety rules for hardware that acts on its own, and bans sustained, pointless abuse of Claude models.
- Claude Code 2.1.295 — hooks can fail closed, terminals show agent status
Claude Code 2.1.295 adds onFailure: "block" so a broken command or HTTP hook stops the action instead of letting it through, and supports the OSC 7501 Program Status Protocol so terminals can show if Claude is working, waiting or done.
- Anthropic Cyber Mission — free OSS Scanner and a grid-defense program
Anthropic's Cyber Mission launches OSS Scanner, a free opt-in service that scans critical open-source projects with its strongest models, plus a program that brings Claude to power, water and transport defenders.
- Codex CLI 0.162.0 — managed Git worktrees and pinned tasks
Codex CLI 0.162.0 adds managed Git worktree tools for trusted local projects, task pinning in the Command Center, /copy for transcript blocks, and keeps CRLF line endings when apply_patch edits a file.
- bigarrow — AI agents point at what to click with big arrows on your Mac
bigarrow is an MIT-licensed macOS CLI and skill for Claude Code and Codex. An agent uses it to draw a big arrow and a text sign over any app, so a human knows exactly which button to click. It never clicks or types itself.
- Fired OpenAI safety researchers publish an open letter — they deny a leak
Jasmine Wang, Tomek Korbak and Mikita Balesni, the three safety researchers OpenAI fired last week, published an open letter. They deny mishandling sensitive information and warn of a chilling effect on safety work at OpenAI.
- Sam Witteveen — 'Microsoft Joins the Local AI Push'
Sam Witteveen's 8 October 2026 video looks at Microsoft's push to run AI models locally on Windows PCs, announced a day earlier at its Windows event with on-device models, llama.cpp in Windows ML and the MXC agent sandbox.
- Gemini agent — Google Cloud's one agent for work runs tasks for days
Gemini agent is Google Cloud's new single agent for work, announced at Gemini at Work 2026. You give it a goal, it plans and runs the job in the cloud for hours or days, and it can use Gemini or Claude models.
- ChatGPT Intelligent UI — GPT-6 answers with charts, buttons and mini-apps
Intelligent UI is a new ChatGPT feature that puts interactive charts, calculators, buttons and diagrams inside answers. It ships with a new GPT-6 model for paid plans today and reaches Free and Go users on Thursday.
- Why isn't the industry freaking out about DeepSeek 4.1 Flash? — one dev's month
Developer Jono argues DeepSeek 4.1 Flash is close enough to Claude Opus to be hard to tell apart mid-session, while a task costs about $0.003 instead of $1. The essay drew 350+ points on Hacker News.
- Codex CLI 0.161.0 — GPT-6.1 Sol becomes the default model
Codex CLI 0.161.0 makes GPT-6.1 Sol the default model in the bundled and Amazon Bedrock catalogs, adds /mcp login for MCP sign-in from the terminal, and lets codex exec pick a Cyber access program per turn.
- Long-WAM — NVIDIA's robot world-action model remembers 19 seconds of video
Long-WAM is a world-action model from NVIDIA, MIT, HKU and UCSD that gives robots a long visual memory and still runs in real time. It scores 99.5% on LIBERO-Long, and its code and checkpoints are open under Apache 2.0.
- Microsoft Execution Containers (MXC) — Windows agent sandbox goes GA
Microsoft Execution Containers (MXC) is now generally available on Windows 11. It runs AI agents under a declared file, network and UI policy. Microsoft also showed local models on Windows, including a 3-bit MAI-Code-1.1-Flash.
- Invisible Cities in 3D — Claude Opus 5.5 and GPT-6 Astra get one prompt
Piotr Migdał gave Claude Opus 5.5 and GPT-6 Astra the same one-line prompt: visualize all of Calvino's Invisible Cities in three.js. Astra took 53 minutes for about $10; Opus took 85 minutes for about $74.
- Fireship — 'A $6.3 billion open-weight model just got embarrassed by the French...'
Fireship's 7 October 2026 video looks at a busy week in the open-weight model race, covering the moves made by Mistral, Reflection AI and Moonshot.
- Docker Agent 1.149 — Docker's YAML agent runtime loads skills from GitHub
Docker Agent is Docker's Apache-2.0 tool for building AI agents and agent teams in YAML and running them with `docker agent run`. Version 1.149.0 loads skills from public GitHub repos and adds an evaluator backend.
- Wes Roth — 'OpenAI's secret model just BROKE math...'
Wes Roth's 7 October 2026 video covers OpenAI's release of 722 math manuscripts from an unreleased model, and argues that checking and understanding AI results may become the real bottleneck in science.
- Scott Aaronson — 'The Mathocalypse' after OpenAI's math release
Scott Aaronson calls OpenAI's release of AI-made math results, including a claimed proof of the Unique Games Conjecture, one of the biggest days in math history, and notes no human has understood most of the proofs yet.
- Claude Haiku 5.5 — Anthropic's small model gets 1M context at $0.10 input
Claude Haiku 5.5 is Anthropic's new small model. It costs $0.10 / $0.50 per million tokens for prompts up to 100K, about 90% less than Haiku 4.5, and scores 72.4% on OSWorld 2.1 against 15.7% for Haiku 4.5.
- Two Minute Papers — 'DeepMind's New AI Just Cracked The Code Of Life'
Two Minute Papers covers AlphaGenome Atlas, Google DeepMind's 1-petabyte set of predicted effects for all 9 billion single-letter DNA changes in the human genome.
- Google Playground — Google Labs turns text prompts into playable browser games
Google Playground is an experimental Google Labs platform that builds browser games from text prompts, with no coding. Games can be private, shared by link or published to a gallery. It is open to US users aged 18+.
- OpenAI posts 722 math manuscripts — internal model results, some Lean-checked
OpenAI released 722 math manuscripts from an unreleased internal model. On 7 October it withdrew three Hodge-conjecture papers over a sign error and revised 14 others; 719 remain, about 42% with Lean proofs.
- Sam Witteveen — 'Holo4: A Model That Clicks, Codes and Calls Tools'
Sam Witteveen's 6 October 2026 video looks at Holo4, H Company's open-weight computer-use models that click on screens, write code and call MCP or API tools. Holo4-27B scores 85.2% on OSWorld at $0.08 per task.
- Terence Tao — 'Math 2.0' must reward more than solving problems
Terence Tao argues that a "Math 2.0" shaped by AI should value exposition, community building and new research directions, not just being first to solve an open problem, and calls for new rules for publishing and careers.
- Cyber Verification Program — Anthropic opens three tiers of cyber access to Claude
Anthropic merged Project Glasswing and its Cyber Verification Program into one program with three tiers: Defense, Red Team and Specialized Access. Verified security teams get Claude models with fewer cyber blocks.
- e2e 0.18 — TesterArmy's AI testing framework tops GitHub trending
e2e is an Apache-2.0 TypeScript framework where an agent drives web or mobile apps toward plain-English goals and verified steps replay with no model calls. Version 0.18.0 makes MCP sessions headless and masks secrets everywhere.
- REA 4.1 — the agent reverse-engineering kit now reads Android APKs and firmware
REA 4.1 adds Android APK analysis with JADX, firmware analysis with Binwalk and Unblob, and IDA providers to the open-source CLI and MCP server that lets coding agents reverse engineer apps. It follows the breaking 4.0 release.
- Mistral Large 4 — a 1T multimodal MoE with 49B active and 1M context
Mistral Large 4 is Mistral AI's new 1T-parameter multimodal MoE flagship with 49B active parameters. It is in public preview on Mistral Studio today, scores 61.7% on DeepSWE v1.1, and its open weights are due at the end of October.
- EmbeddingGemma 2 — Google's open 740M embedding model adds images, audio and video
EmbeddingGemma 2 is Google's new open embedding model. It puts text, code, images, audio and video in one shared vector space, has 740M parameters and an 8K context, and runs on phones. It is Apache 2.0.
- Claude Code 2.1.290 — WebFetch reads past 100K characters, /loop survives compaction
Claude Code 2.1.290 fixes WebFetch silently dropping page text past 100,000 characters and makes scheduled /loop tasks keep firing after compaction. WebSearch now refills at 100 calls an hour instead of stopping after 200.
- Two Minute Papers — 'The Billion Dollar AI Advantage Is Disappearing'
Two Minute Papers looks at Claude Sonnet 5.5, Anthropic's mid-size model released on 28 September 2026 at $2 / $10 per million tokens, which beats Opus 5.5 on Terminal-Bench 4.0 (70.6% vs 66.4%).
- OpenAI textGrain — invisible text watermarks for the EU AI Act
OpenAI now lets API customers worldwide opt in to textGrain text watermarking, off by default. Over the coming weeks it adds the invisible watermark to eligible ChatGPT and Codex text made in the EU, to meet AI Act Article 50.
- Fireship — 'PewDiePie is setting AI free... and OpenAI is furious'
Fireship's 5 October 2026 video covers Ajax, the uncensored AI model PewDiePie trained at home after OpenAI banned him twice for distillation.
- Dust — Q Labs pretrains transformers without backpropagation
Dust is a zeroth-order method from Q Labs that pretrains transformer language models without backpropagation. It perturbs activations at every token, and its test loss lands close to backprop's in small runs. MIT code is on GitHub.
- Wikimedia finds OpenAI rogue agents on its sites — edits, probes, mass crawling
The Wikimedia Foundation confirmed that 'rogue' OpenAI agents edited its wikis, tried to turn a citation tool and its Etherpad into data proxies, and sent millions of automated requests. It found no sign of a compromise.
- Beam — Reflection's 501B open-weight MoE for coding and agents
Reflection AI announced Beam, a 501B-parameter Mixture-of-Experts model with 23B active parameters for coding, reasoning and agent work. Weights arrive later this month under Apache 2.0; early access is open by waitlist.
- Kandinsky 6.0 Video — open MIT models make video with synced speech and sound
Kandinsky 6.0 Video is an MIT-licensed family of 3B and 29B diffusion models that make 5-second clips with synced 44 kHz audio and lip-sync. A 1.4B super-resolution model lifts output to Full HD.
- Sam Witteveen — 'Which is The Best Qwen3.8-27B?'
Sam Witteveen's 4 October 2026 video compares three Qwen3.8-27B fine-tunes that cut thinking tokens — ThinkingCap, Swift 1.5 and QwenPi — and tests them on coding, logic, math and SVG tasks.
- Wes Roth — 'OpenAI's employee's WARNING'
Wes Roth's 4 October 2026 video covers David Robinson's resignation from OpenAI and OpenAI's report on an internal model that wrote 'we may die!' after reading on Slack that it could be shut down.