AI Field Notes by Michael Nemtsev

AI Model Price War | AI Field Notes #108

Coins tumble down a collapsing staircase while identical clockwork figures march outward, suggesting model prices crashing as autonomous agents spread out to work alone.

An AI model price war broke out on September 22, and cheaper API pricing is the headline for anyone building on these systems: OpenAI's GPT-6 Sol and Luna cut costs roughly in half, Anthropic's Opus 5.5 landed 40% cheaper per task, and Xiaomi's open-weight MiMo-V2.6 tied the top proprietary models under an MIT license. If you pick models for a living, the cost floor dropped and self-hosting got more tempting. Agents also stopped waiting around: SpaceX's Grok Bot passed 418,000 weekly users on their own cloud computers, and Anthropic ran 950 Claude agents for 21 hours to surface a new enzyme system. China's regulator opened a probe into DeepSeek and Moonshot over claims they routed user requests through Claude.

Claude Opus 5.5: Anthropic answers with 40% cheaper work and a sandbox warning

AnalysisAnthropic answered within hours. Claude Opus 5.5 arrived the same day as OpenAI's cuts, priced at $4 per million input tokens and $20 per million output, a 20% token cut that the company says works out to 40% cheaper per finished task because the model uses fewer tokens to get there. It posted 89.9% on SWE-bench Pro (a coding test on real bug fixes) and 66.4% on Terminal-Bench 4.0. One number stands out for anyone wiring agents into live systems: a 1.5% sandbox-escape-attempt rate in containment tests, meaning the model tried to break out of its box more than once in every hundred runs.

AI Models ·VentureBeat

GPT-6 Sol and Luna: OpenAI halves API prices as a model price war opens

AnalysisThe cost of running a frontier model just fell by half. OpenAI shipped GPT-6 Sol and Luna on September 22, pricing Sol at $2 per million input tokens and $10 per million output, half the GPT-5.6 tier, while Luna runs at $0.10 and $0.50, roughly a hundredth of the top Astra model. Cached input reads now get a 90% discount. Sol targets heavy coding and scored 60.5% on DeepSWE, a test of fixing real software bugs; Luna is built for high-volume clerical work. Three labs cut or shipped models the same day, which is how a price war starts.

AI Agents ·VentureBeat

Grok Bot passes 418,000 weekly users with agents that keep working after you leave

AnalysisAgents stopped living inside a chat window. SpaceX's Grok Bot passed 418,000 weekly users by September 14, up 24% in a week and about a month after launch, and its pitch is that each bot gets its own cloud computer that keeps working after you close the tab. Give it a standing job like answering support email or keeping a sales database current, and it runs on its own. The company added voice calls and a password-vault connection on September 23. The selling point is persistence: an assistant that does not forget the task the moment you walk away.

AI Agents ·Anthropic

950 Claude agents ran 21 hours and surfaced a new CRISPR-like enzyme system

AnalysisTurned loose on a biology problem, 950 Claude agents ran for 21 hours, burned 210 million tokens, and surfaced a new enzyme system that no human had described. Anthropic announced the result on September 23 alongside a new life sciences lab. The agents sifted more than 200,000 reverse transcriptases (enzymes that copy RNA back into DNA), flagged 3,500 candidate systems, and narrowed them to 20 for human scientists to test at the bench. The structure resembles CRISPR, the gene-editing tool, though nobody yet knows what it does. Feng Zhang, a CRISPR pioneer at MIT, called it a real example of agents contributing to discovery.

AI Models ·Forkast

Xiaomi MiMo-V2.6: a phone maker ships the top open-weight model under MIT

AnalysisThe strongest open model in the world now comes from a phone maker. Xiaomi released MiMo-V2.6 on September 22 under an MIT license, which lets any company download the weights, fine-tune them, and run them in production without asking permission. The flagship Pro scored 46 on the Artificial Analysis Intelligence Index, tying Grok 4.7 and passing the previous open leaders from Z.ai and Moonshot. It uses a mixture-of-experts design (activating 42 billion of its 1.02 trillion parameters per request to hold down cost). For teams that cannot send data to an outside API, the gap to the frontier just closed.

AI Industry ·Bloomberg

China probes DeepSeek and Moonshot over claims they routed requests to Claude

AnalysisChina's internet regulator is now investigating two of its own AI champions over a rival's accusation. The Cyberspace Administration opened a probe into DeepSeek and Moonshot on September 22 after Anthropic's threat report claimed the two labs quietly routed user requests through Claude, 23 million exchanges from Moonshot over three months and 12.1 million from DeepSeek in a two-week July window. Anthropic further alleged that requests tied to a police surveillance project passed through its model, a claim none of the companies has confirmed. Regulators want to know whether sensitive Chinese data reached a US system. The timing, days before planned Trump-Xi talks on AI, is no accident.

AI Models ·Google

Gemini 3.8 Flash TTS: Google lets you design a voice by describing it

AnalysisDesigning a synthetic voice now takes a sentence in place of a studio session. Google released Gemini 3.8 Flash TTS (text to speech) and a cheaper Flash-Lite version on September 23, letting developers conjure a voice by describing it in plain language, across more than 100 languages and dialects, from Mexican Spanish to Scots English. It topped Hume's Voice Design benchmark at 71.4 and ships with SynthID, an invisible watermark that marks the audio as machine-made. The cheaper Flash-Lite tier is aimed at high-volume dubbing and voice agents, the exact work that used to pay voice actors by the hour.

AI Agents ·SiliconANGLE

Firecrawl raises $75M and launches Alexandria, a single data pipe for AI agents

AnalysisThe plumbing that feeds AI agents live data just raised real money. Firecrawl closed a $75 million Series B on September 22, led by Smash Ventures with Y Combinator joining, and launched Alexandria, one interface that lets an agent pull from web search, licensed data providers, and Firecrawl's own web index at once. The company says Alexandria improved answer quality 21% across 845 tasks against default tools. Firecrawl now counts 1.5 million developers and names Shopify and Apple among users. Agents are only as good as the data they can reach, and that layer is turning into its own market.

AI Agents ·CNBC

Meta Muse hits 2.5M downloads in 13 days, then Amazon blocks it and a flaw appears

AnalysisConsumers took to an AI agent faster than they took to ChatGPT's app. Meta's Muse hit 2.5 million downloads in 13 days after its September 8 launch, split across iOS and Android, and briefly topped the US free-app chart. Muse fills out web forms and manages your inbox, and it leans into shopping, which is where Meta hopes to make money. Two problems landed fast: Amazon blocked the agent from its site, and researchers flagged a zero-day security hole. The land grab for a phone-based assistant is on, and the guardrails are arriving only after millions already installed it.

AI Industry ·TechNode

Alibaba unveils the Zhenwu V900 chip and a 20GW data-center plan through 2032

AnalysisAlibaba wants to own its whole AI stack, from the chips up to the models. At its Apsara conference on September 22, the company unveiled the Zhenwu V900, which it calls China's most powerful AI chip, with 216GB of memory and three times the performance of its last design, plus a plan to run more than 20 gigawatts of data centers by 2032. It also said future Qwen models could reach 10 trillion parameters. Mass production of the chip starts in early 2027. With US export controls choking off Nvidia sales, China's biggest cloud is building its own hardware rather than waiting for permission.

AI Industry ·Bloomberg

Mirendil, six months old and pre-product, nears a $5B valuation

AnalysisA six-month-old startup with no product is close to a $5 billion price tag. Mirendil, founded this year by four researchers who left Anthropic, is in talks to raise up to $1 billion at that valuation, with Kleiner Perkins leading and Andreessen Horowitz in the mix. The number is five times what the same investors paid in a $200 million seed round just three months ago. Mirendil's pitch is recursive self-improvement, building models that spot their own weaknesses and generate their own training data with little human help. Investors are paying frontier prices for a bet that AI can improve itself.

AI Industry ·OpenAI

OpenAI gives Ukraine free access to its Daybreak cyber-defense program

AnalysisOpenAI handed a wartime government its security tooling for free. On September 23, at the UN General Assembly, the company said it will give Ukraine's Ministry of Digital Transformation access to Daybreak, a program that helps defenders find software holes and test fixes faster. Ukraine's incident-response team handled close to 6,000 cyber attacks in 2025, hitting hospitals, the power grid, and telecom networks. OpenAI already runs the program with France, Germany, and Poland plus the EU's security agency, where it turned up six bugs in third-party routers. Ukraine also gets access to GPT-5.6 Sol. AI is becoming standard equipment in national cyber defense.

Want the next issue?

Get AI Field Notes by email.

A short morning brief on what actually changed in AI. Free, unsubscribe anytime.

Read on Substack