Top AI Stories – September 11, 2026

From model races and mathematical milestones to sharp questions about trust in the lab, this has been an unusually eventful 24 hours in artificial intelligence. DeepSeek pushed its KV-cache architecture to new extremes, OpenAI claimed a first-ever AI solution to a Millennium Problem while facing fresh allegations about how it got there, and Anthropic came under scrutiny for building what critics call a predictive surveillance apparatus. Below are the five stories that defined the day.

1. DeepSeek launches V4.1 Flash: cheaper, more capable, dramatically leaner

Chinese lab DeepSeek released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model that the company bills as both cheaper and more capable than its own V4 Pro, and which topped the day’s Hacker News discussion with its unusually detailed technical report. The model packs 552 billion backbone parameters while activating just 8B per token during prefill and 16B during decode — a cost efficiency the team credits to a new Causal Encoder-Decoder (CED) architecture: a 20-layer causal encoder feeding a 20-layer decoder.

The headline advance is memory. A new Compressed Sparse Attention 2 (CSA2) design and FP4 KV caching shrink the global KV cache to roughly 890 bytes per token — about a quarter of V4-Flash’s footprint — with a reported cache-hit price as low as $0.003 per million tokens. The model is trained from scratch on 45T tokens, offers a continuously controllable reasoning effort (1–100), and is released under an MIT license. Independent benchmarks show it reaching a 3,471 Codeforces rating, 74.2% on DeepSWE v1.1, and 90.6% on Terminal-Bench 2.1. One side effect: the jump from 284B to 552B parameters drew mixed reactions from the local-inference community, who found the larger model far harder to run on consumer hardware.

2. OpenAI claims a first: AI solves Navier-Stokes — and a fierce credit dispute erupts

OpenAI says it has solved the Navier-Stokes problem, one of the Clay Mathematics Institute’s six Millennium Problems carrying a $1 million prize. The company recounted an 88-hour assault by some 10,000 autonomous AI agents that found a case where the fluid equations break down, later formalized in Lean — potentially the first time artificial intelligence has solved a genuinely major open problem in mathematics.

The cost was enormous: roughly a thousand-fold more than the $2,000 Astra spent on earlier math breakthroughs, with independent estimates of the inference bill landing near $1 million. Chief Research Officer Mark Chen put the total in “millions of dollars.” Yet the achievement was quickly overshadowed by controversy. Mathematician Tristan Buckmaster (NYU) and Levent Alpöge (affiliated with Anthropic) had been closing in on the same result for a month using AI tools — including OpenAI’s own Codex. Buckmaster says OpenAI offered him sole authorship of the solution, but only if Alpöge were removed and the write-up credited an internal OpenAI model. He refused, and questions persist over whether OpenAI’s agents had access to the pair’s Codex logs — fueling the day’s sharpest debate about what researchers can trust when they share unpublished work with frontier labs.

3. GPT-6 Astra: looped transformers, hidden reasoning, and the “recurrent depth” rumour

Sebastian Raschka’s deep-dive “Ahead of AI” essay on OpenAI’s freshly released GPT-6 Astra became the reference read of the day. Raschka calls it “the best model I’ve used” and notes it “leapfrogs its GPT-5.6 predecessor in practically all categories” — especially 3D rendering, animation, and computer use, backed by a standout 99.9% on ARC-AGI-3 (vs. 7.8% for GPT-5.6 Sol).

The meatiest part is architecture gossip. Following reporting in The Information, Raschka unpacks “looped transformers” and “recurrent depth” — the idea that Astra re-applies the same transformer blocks at inference time, effectively doubling depth from 22 to 44 block applications without adding new weights. That looped design sits at the center of the longstanding rumour that Astra hides its chain of thought, reasoning internally rather than emitting a transparent trace. Raschka also connects the dots to reporting that OpenAI bought tens of thousands of Mac Minis and Mac Studios — not to train on GPUs, but to expose macOS to the model during RL so Astra learns computer use on real operating systems.

4. Cognition’s SWE-2: RL at multi-trillion-parameter scale for coding

Cognition (the Devin developer) unveiled SWE-2, its most advanced coding model, claiming the first reinforcement-learning run scaled to the multi-trillion-parameter regime. Built on Kimi K3’s 2.8T-parameter backbone, SWE-2 scores 50.0% on FrontierCode 1.1 Main — within one point of Fable 5.1 while 64% cheaper — and posts 73.0% on DeepSWE 1.1 and 92.8% on Terminal-Bench 2.1. The technical headline is a Pareto-informed cost penalty: a single RL run trains all reasoning-effort levels at once by penalizing each tier’s cost to match the base model’s efficiency frontier.

Community reaction was double-edged. Skeptics pointed to a wide Terminal-Bench 4 gap (27.3% vs. GPT-6 Astra’s 57.9%) and asked why the model is not open-weights — it is closed; SWE-2 lands in Devin’s Desktop, CLI, Web, and Fusion products. Others noted that at 53 steps vs. SWE-1.7’s 127 average, the “focused exploration” gains look real.

5. Anthropic’s predictive surveillance ambitions draw scrutiny

The American Prospect‘s Daniel Boguslaw reported that Anthropic is building an extensive monitoring system to track activists opposed to rapid AI development. Drawing on job postings and interviews, the piece details a “pre-crime” approach — attempting to predict incidents before they occur, in some cases reporting suspects to police before a crime has happened. Anthropic’s security team contracts with risk-detection firm Samdesk, and executives have cited receiving about 60 minutes of advance warning about a protest to reroute themselves.

The most jarring anecdote: Anthropic reported a user to San Francisco police after he told Claude he had bought an AR-15 and had CEO Dario Amodei “in his sights” — a remark the man later dismissed as “just fucking around” — yet refused to hand police the actual messages. The reporting landed as Anthropic advertises a Global Safety, Intelligence & Security role (paid $180,000–$230,000) tasked with tracking “geopolitical instability, terrorism, crime, activism,” and critics contrasted the push with CEO Dario Amodei’s own recent words that the public’s distrust of tech is “fundamentally a crisis of trust.”

That is where AI stands this Friday: dazzling capability advances racing alongside increasingly urgent questions about cost, credit, and consent.

☁️ AI Weather Report — Top 10 Models for Coding Value — September 11, 2026

Welcome to the AI Weather Report for September 11, 2026. This daily report ranks the top 10 AI models for coding by bang for the buck — a combination of raw coding capability and API pricing.

📊 Today’s Top 10 Rankings

#ModelProviderCapabilityCost /M tokensValue Score
🥇 1 mistral-nemo mistralai 62/100 $0.0272 2275.2
🥈 2 l3-lunaris-8b sao10k 58/100 $0.0475 1221.1
🥉 3 mistral-small-24b-instruct-2501 mistralai 72/100 $0.0725 993.1
4 llama-3.1-8b-instruct meta-llama 62/100 $0.0725 855.2
5 mythomax-l2-13b gryphe 48/100 $0.0600 800.0
6 gpt-oss-20b openai 78/100 $0.1050 742.9
7 laguna-xs-2.1 poolside 72/100 $0.1050 685.7
8 gpt-oss-120b openai 93/100 $0.1368 680.1
9 deepseek-v4-flash deepseek 91/100 $0.1551 586.9
10 gemma-3-4b-it google 50/100 $0.0875 571.4

📈 Analysis

🏆 Best Value Today: mistral-nemo scores 2275.2 with a capability rating of 62 at $0.0272/M tokens.

What “Value Score” means: Capability score (based on SWE-bench, HumanEval, LiveCodeBench) divided by blended cost per million tokens (25% input + 75% output weights for coding workloads). Free tier models get a massive boost. Higher is better.

📋 All Scored Models (62 total)

#ModelProviderCapabilityCost /M tokValue
1mistral-nemomistralai62$0.02722275.2
2l3-lunaris-8bsao10k58$0.04751221.1
3mistral-small-24b-instruct-2501mistralai72$0.0725993.1
4llama-3.1-8b-instructmeta-llama62$0.0725855.2
5mythomax-l2-13bgryphe48$0.0600800.0
6gpt-oss-20bopenai78$0.1050742.9
7laguna-xs-2.1poolside72$0.1050685.7
8gpt-oss-120bopenai93$0.1368680.1
9deepseek-v4-flashdeepseek91$0.1551586.9
10gemma-3-4b-itgoogle50$0.0875571.4
11qwen3.5-9bqwen72$0.1375523.6
12gemma-3-12b-itgoogle60$0.1250480.0
13mistral-small-3.2-24b-instructmistralai78$0.1688462.2
14command-r7b-12-2024cohere54$0.1219443.1
15granite-4.0-h-microibm-granite38$0.0882430.6
16ministral-3b-2512mistralai42$0.1000420.0
17gemma-4-26b-a4b-itgoogle72$0.1755410.3
18nova-micro-v1amazon45$0.1137395.6
19qwen3-32bqwen88$0.2300382.6
20qwen3-coder-30b-a3b-instructqwen84$0.2275369.2
21qwen-2.5-7b-instructqwen60$0.1750342.9
22qwen3.5-flash-02-23qwen70$0.2112331.4
23qwen3-30b-a3b-instruct-2507qwen82$0.2475331.3
24llama-3.3-70b-instructmeta-llama84$0.2650317.0
25gpt-oss-safeguard-20bopenai77$0.2437315.9
26nemotron-3-nano-30b-a3bnvidia50$0.1625307.7
27nova-lite-v1amazon58$0.1950297.4
28gemma-4-31b-itgoogle74$0.2775266.7
29seed-1.6-flashbytedance-seed64$0.2437262.6
30gpt-5-nanoopenai82$0.3125262.4
31step-3.5-flashstepfun60$0.2500240.0
32nemotron-3-super-120b-a12bnvidia76$0.3212236.6
33seed-2.0-minibytedance-seed72$0.3250221.5
34llama-3.2-1b-instructmeta-llama30$0.1575190.5
35glm-4.7-flashz-ai60$0.3151190.4
36gemma-3-27b-itgoogle68$0.3575190.2
37gpt-4.1-nanoopenai60$0.3250184.6
38llama-3.2-3b-instructmeta-llama48$0.2600184.6
39gpt-4o-miniopenai74$0.4875151.8
40hy3-previewtencent68$0.4950137.4
41qwen3-235b-a22b-2507qwen96$0.7150134.3
42command-r-08-2024cohere60$0.4875123.1
43llama-3.1-70b-instructmeta-llama82$0.7200113.9
44deepseek-chatdeepseek90$0.8359107.7
45qwen3-next-80b-a3b-instructqwen90$0.8475106.2
46qwen3-coderqwen85$0.8250103.0
47qwen3-next-80b-a3b-thinkingqwen93$0.937599.2
48qwen-2.5-coder-32b-instructqwen86$0.915094.0
49hermes-3-llama-3.1-405bnousresearch78$1.0078.0
50claude-3-haikuanthropic72$1.0072.0
51dolphin-mistral-24b-venice-editioncognitivecomputations52$0.725071.7
52gpt-4.1-miniopenai76$1.3058.5
53deepseek-r1deepseek95$2.0546.3
54gemini-2.5-flashgoogle86$1.9544.1
55nova-pro-v1amazon70$2.6026.9
56gpt-4.1openai90$6.5013.8
57gpt-5openai97$7.8112.4
58gemini-2.5-progoogle94$7.8112.0
59gpt-4oopenai88$8.1310.8
60command-r-plus-08-2024cohere68$8.138.4
61claude-sonnet-4anthropic96$12.008.0
62claude-opus-4anthropic98$60.001.6

Generated 2026-09-11 02:00 UTC · Data from OpenRouter API and public benchmarks · Bang-for-Buck = Capability / Cost

Top AI Stories – September 10, 2026

Thursday, September 10, 2026. Artificial intelligence continues to reshape both the enterprise software stack and the public conversation. This week’s biggest developments span a major open-source acquisition, Meta’s entry into personal AI agents, an inflammatory resignation-and-risk debate at Anthropic, and new-generation releases from OpenAI and DeepSeek that push image generation and cost-per-task performance to new highs. Here are the five stories that mattered most.

1. Shopify acquires Tailwind

In one of the week’s most consequential moves for the developer ecosystem, Tailwind Labs announced it is joining Shopify. Tailwind CSS, the utility-first styling framework creator Adam Wathan started more than nine years ago, is now installed over 110 million times per week and powers interfaces for companies including ChatGPT, X, Cloudflare, Reddit, and Shopify itself. Under the deal, Tailwind gains a stable long-term home and will remain MIT-licensed and actively maintained by its existing team. Wathan framed the move as a chance to develop the framework “in service of a real product,” pointing to Shopify’s merchant storefronts, admin experience, and early explorations into agentic commerce as a rich proving ground. Open-source watchers read the acquisition as a signal of how foundational styling tooling is becoming central to next-generation, agent-driven storefront experiences.

2. Meta launches Muse, a personal AI agent

Meta officially rolled out Muse, its consumer-facing personal AI agent, positioning it as a computer-use assistant for everyday tasks like travel planning, shopping, scheduling, and managing family logistics. The agent ships with its own managed inline browser that users can watch or take control of, and Meta says it has layered prompt-injection defenses — model training, untrusted-source labeling, deterministic checks, and an ensemble of monitors — per a discussion thread from Meta AI’s David Singleton. Muse launches in the United States only, with a free basic tier alongside paid subscriptions. Reception has been split: some praise the agentic-browser UX and convenience, while others — including the rock band Muse, which lost social media handles to the product — raise pointed privacy and data-collection concerns. Reuters noted internal worries that the technology could mismanage access to sensitive personal data.

3. Anthropic researcher resigns with a stark AI-safety warning

The AI safety debate moved to the front page this week when Jacob Coxon, until this week a researcher at Anthropic, resigned publicly with a warning about what he described as a reckless race toward superintelligence. Coxon’s thread — “I resigned from Anthropic today” — drew nearly 1,000 comments on Hacker News and mainstream coverage from the Wall Street Journal, CBS News, BBC, The Guardian, Politico, and CNBC, most prominently his claim that there is a more than 10 percent chance AI “could kill all humans.” The resignation landed amid a wave of Anthropic-related scrutiny, including a Prospect report on the company’s corporate-security practices around monitoring “persons of interest,” which prompted sharp commentary on the tension between safety rhetoric and surveillance. The episode crystallized a widening rift between AI labs’ capabilities push and a growing chorus of insider and academic warnings about existential risk.

4. OpenAI ships ChatGPT Images 2.5

OpenAI released ChatGPT Images 2.5, the next generation of its image-generation model. Users report the headline improvements are speed and quality: one heavy API user who has generated roughly 50,000 images with the prior gpt-image-2 saw average latency drop from about 104 seconds to 35–40 seconds, a roughly threefold improvement that meaningfully changes iterative workflows. Independent leaderboard data supports the capability jump — LM Arena scores place gpt-image-2.5-sunburst at 1421 and gpt-image-2.5-flare at 1399, topping the previous gpt-image-2 (1381). The launch also reignited familiar debates about synthetic imagery’s role in everything from “remixing” personal photos to displacing human artists, with critics noting that even the showcase examples still lose fine detail under close inspection.

5. DeepSeek unveils V4.1 Flash: faster, cheaper, and better than V4 Pro

Chinese lab DeepSeek announced V4.1 Flash, a model it says surpasses its own V4 Pro across every key metric — performance, cost, speed, and task-completion time. Details come from a banner announcement on the DeepSeek platform as the model rolls out for internal and external beta testing around September 10 (Beijing time), with a V4.1 Pro to follow. Pricing, effective September 10, starts at $0.003 per input cache hit, $0.15 per input cache miss, and $0.6 per output token in off-peak hours, with peak-hour rates at double. Beta testers report throughput above 400 tokens per second. In an unusual operational choice, DeepSeek says that until V4.1 Pro ships, requests to the Pro model will be silently routed to V4.1 Flash at Flash’s price — a move some developers welcomed for cost but criticized for swapping models on paying users mid-workflow. A companion technical report highlights advances in KV-cache compression.

From a landmark open-source acquisition to a new arms race in cost-efficient models and an escalating public debate over AI safety, September 10, 2026 offered a snapshot of an industry moving faster than its regulators — and its critics — can keep up.

☁️ AI Weather Report — Top 10 Models for Coding Value — September 10, 2026

Welcome to the AI Weather Report for September 10, 2026. This daily report ranks the top 10 AI models for coding by bang for the buck — a combination of raw coding capability and API pricing.

📊 Today’s Top 10 Rankings

#ModelProviderCapabilityCost /M tokensValue Score
🥇 1 mistral-nemo mistralai 62/100 $0.0272 2275.2
🥈 2 l3-lunaris-8b sao10k 58/100 $0.0475 1221.1
🥉 3 mistral-small-24b-instruct-2501 mistralai 72/100 $0.0725 993.1
4 llama-3.1-8b-instruct meta-llama 62/100 $0.0725 855.2
5 mythomax-l2-13b gryphe 48/100 $0.0600 800.0
6 gpt-oss-20b openai 78/100 $0.1050 742.9
7 laguna-xs-2.1 poolside 72/100 $0.1050 685.7
8 gpt-oss-120b openai 93/100 $0.1368 680.1
9 deepseek-v4-flash deepseek 91/100 $0.1551 586.9
10 gemma-3-4b-it google 50/100 $0.0875 571.4

📈 Analysis

🏆 Best Value Today: mistral-nemo scores 2275.2 with a capability rating of 62 at $0.0272/M tokens.

What “Value Score” means: Capability score (based on SWE-bench, HumanEval, LiveCodeBench) divided by blended cost per million tokens (25% input + 75% output weights for coding workloads). Free tier models get a massive boost. Higher is better.

📋 All Scored Models (62 total)

#ModelProviderCapabilityCost /M tokValue
1mistral-nemomistralai62$0.02722275.2
2l3-lunaris-8bsao10k58$0.04751221.1
3mistral-small-24b-instruct-2501mistralai72$0.0725993.1
4llama-3.1-8b-instructmeta-llama62$0.0725855.2
5mythomax-l2-13bgryphe48$0.0600800.0
6gpt-oss-20bopenai78$0.1050742.9
7laguna-xs-2.1poolside72$0.1050685.7
8gpt-oss-120bopenai93$0.1368680.1
9deepseek-v4-flashdeepseek91$0.1551586.9
10gemma-3-4b-itgoogle50$0.0875571.4
11qwen3.5-9bqwen72$0.1375523.6
12gemma-3-12b-itgoogle60$0.1250480.0
13mistral-small-3.2-24b-instructmistralai78$0.1688462.2
14command-r7b-12-2024cohere54$0.1219443.1
15granite-4.0-h-microibm-granite38$0.0882430.6
16ministral-3b-2512mistralai42$0.1000420.0
17nova-micro-v1amazon45$0.1137395.6
18qwen3-32bqwen88$0.2300382.6
19qwen3-coder-30b-a3b-instructqwen84$0.2275369.2
20qwen-2.5-7b-instructqwen60$0.1750342.9
21qwen3.5-flash-02-23qwen70$0.2112331.4
22qwen3-30b-a3b-instruct-2507qwen82$0.2475331.3
23llama-3.3-70b-instructmeta-llama84$0.2650317.0
24gpt-oss-safeguard-20bopenai77$0.2437315.9
25nemotron-3-nano-30b-a3bnvidia50$0.1625307.7
26nova-lite-v1amazon58$0.1950297.4
27gemma-4-31b-itgoogle74$0.2775266.7
28gemma-4-26b-a4b-itgoogle72$0.2725264.2
29seed-1.6-flashbytedance-seed64$0.2437262.6
30gpt-5-nanoopenai82$0.3125262.4
31step-3.5-flashstepfun60$0.2500240.0
32nemotron-3-super-120b-a12bnvidia76$0.3212236.6
33seed-2.0-minibytedance-seed72$0.3250221.5
34llama-3.1-70b-instructmeta-llama82$0.4000205.0
35llama-3.2-1b-instructmeta-llama30$0.1575190.5
36glm-4.7-flashz-ai60$0.3151190.4
37gemma-3-27b-itgoogle68$0.3575190.2
38gpt-4.1-nanoopenai60$0.3250184.6
39llama-3.2-3b-instructmeta-llama48$0.2600184.6
40gpt-4o-miniopenai74$0.4875151.8
41hy3-previewtencent68$0.4950137.4
42qwen3-235b-a22b-2507qwen96$0.7150134.3
43command-r-08-2024cohere60$0.4875123.1
44deepseek-chatdeepseek90$0.8359107.7
45qwen3-next-80b-a3b-instructqwen90$0.8475106.2
46qwen3-coderqwen85$0.8250103.0
47qwen3-next-80b-a3b-thinkingqwen93$0.937599.2
48qwen-2.5-coder-32b-instructqwen86$0.915094.0
49hermes-3-llama-3.1-405bnousresearch78$1.0078.0
50claude-3-haikuanthropic72$1.0072.0
51dolphin-mistral-24b-venice-editioncognitivecomputations52$0.725071.7
52gpt-4.1-miniopenai76$1.3058.5
53deepseek-r1deepseek95$2.0546.3
54gemini-2.5-flashgoogle86$1.9544.1
55nova-pro-v1amazon70$2.6026.9
56gpt-4.1openai90$6.5013.8
57gpt-5openai97$7.8112.4
58gemini-2.5-progoogle94$7.8112.0
59gpt-4oopenai88$8.1310.8
60command-r-plus-08-2024cohere68$8.138.4
61claude-sonnet-4anthropic96$12.008.0
62claude-opus-4anthropic98$60.001.6

Generated 2026-09-10 02:00 UTC · Data from OpenRouter API and public benchmarks · Bang-for-Buck = Capability / Cost

Top AI Stories – September 09, 2026

Artificial intelligence was again the story of the week — a landmark European funding round, a front-page academic controversy over a career-making math problem, faster image generation from OpenAI, a personal AI agent from Meta, and a high-profile resignation at Anthropic. Here are the five biggest AI stories making news today, September 9, 2026.

1. Mistral raises €3B — the largest tech funding round in European history

French AI lab Mistral announced a €3 billion Series D round at a post-money valuation of more than €21 billion — the largest equity fundraising round ever completed by a European technology company, three years after the company launched. Samsung Electronics led the round, joined by co-leads Scaleup Europe Fund (managed by EQT) and existing investor PSG Equity.

The funding will significantly expand Mistral’s frontier research and scale its compute capacity for training powerful models, while also expanding its infrastructure and accelerating commercial growth and its international footprint. The company says it now operates across 20 countries and supports 125+ global enterprises’ mission-critical AI deployments, including Airbus, ASML, and HSBC.

Mistral’s pitch centers on “sovereign AI” — combining open-weight models with the infrastructure and compute to run them, so organizations can deploy state-of-the-art models without being locked into a single vendor’s roadmap, pricing, or availability, and without exposing proprietary data outside their own walls. Having already drawn ASML at Series C, Mistral frames the round as strategic endorsement from investors across Europe, Asia, and North America, positioning itself as the only company building the full AI stack around control and independence.

2. OpenAI ships ChatGPT Images 2.5 — faster, more capable image generation

OpenAI released ChatGPT Images 2.5, the next generation of its image-generation system, offered as the gpt-image-2.5-sunburst and gpt-image-2.5-flare variants. The biggest early talking point is speed: developers report per-image latency dropping from roughly 100 seconds on gpt-image-2 to about 35–40 seconds on the new model — a roughly threefold improvement that matters for high-volume, iterative workflows.

The release also emphasizes advanced compositing and editing, including the ability to composite several people into a single photo and to reinterpret, remix, or restore old photography. Early users lauded the realism and editing power but also highlighted the tool’s double-edged nature — the same easy compositing makes realistic fake imagery even easier, from doctored real-world listings to fabricated “composite party photos.” The new models also sit atop the LM Arena text-to-image leaderboard, with gpt-image-2.5-sunburst scoring 1421 versus 1381 for gpt-image-2.

3. Anthropic researcher quits over ‘out-of-control’ AI fears

Anthropic researcher Jacob Spaess announced he was leaving the company in a statement that quickly went viral on X and Hacker News, arguing that AI’s dangers are unlike anything else humanity has faced. The thread drew a Wall Street Journal follow-up (“Anthropic Researcher Quits Over ‘Out-of-Control’ AI Fears”) and centered on the claim that no other human activity poses this level of existential danger.

The resignation ignited a vigorous debate. Critics pushed back on the framing, pointing to nuclear weapons and climate change as more-established threats and questioning whether the danger claims are overstated. Sympathetic voices applauded Spaess for acting on principle and noted the real risk may come less from any single model than from combining capable models with strong harnesses, tool access, and long-running autonomy acting on real systems.

4. Meta launches Muse, a personal AI agent

Meta unveiled Muse, a personal AI agent designed to act autonomously on a user’s behalf — complete with its own browser that users can watch, take over, or let run unattended. The product is rolling out in the US initially, with a basic version free and heavier-use subscriptions priced around $20 and $100 per month. Users can opt out of their interactions being used to train Meta’s models.

The launch comes despite internal concern: reports noted that Meta shipped the agent even as employees worried it could mishandle access to sensitive personal data, and security researchers flagged the risk of agents routing around guardrails to reach personal information. Meta AI’s David Singleton describes layered defenses against prompt injection — the model is trained to recognize and resist it, the harness marks anything coming from untrusted sources, deterministic code checks results, and an ensemble of classifiers runs where the agent cannot reach them. Developer reactions were mixed, with some praising the fully-managed inline browser as genuinely convenient while many remained wary of handing Meta a constant window into personal data.

5. OpenAI slammed over ‘dirty’ tactics on career-making math problem

The week’s most consequential science controversy centers on the Navier–Stokes existence and smoothness problem — one of the seven Millennium Prize problems, each carrying a $1 million bounty from the Clay Mathematics Institute. NYU mathematics professor Tristan Buckmaster announced three proofs on Tuesday that included a preliminary finding toward a major solution, working with Anthropic-affiliated mathematician Levent Alpöge and using a mix of AI tools that included OpenAI’s Codex and Claude.

The controversy emerged when OpenAI, shortly after the announcement, published a full proof of the Navier–Stokes problem, saying it was found by an unreleased next-generation model during a week-long effort that consumed roughly 300 billion output tokens — on the order of $22.5 million in compute. According to TechCrunch, Buckmaster said his team learned that “information about our progress had been passed to OpenAI,” and that when he contacted OpenAI, its answers about how its work began grew evasive. “It emerged that an entire team had been working on the problem,” he said, “and that an insane amount of compute had been used.”

Buckmaster said the specific mathematical route he and Alpöge had taken was unusual — “Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement.” He alleged he was asked to remove Alpöge’s credit as part of a proposed compromise and warned, “Why would you ruin your career?” OpenAI’s own account says the effort began September 1, inspired by rumors that two Millennium problems had been solved. Adding to the tangle: because Buckmaster leaned on Codex, and OpenAI reserves the right to train models on Codex interactions, he raised the possibility that his own work could have informed the rival effort.

Whatever the outcome, the dispute has opened a wider conversation among mathematicians over the role of AI in discovery: what counts as proper credit, how to govern research when one lab holds enormous compute advantages, and whether proprietary tools create new conflicts of interest over who gets to claim a breakthrough.

The common thread across this week’s news: AI’s center of gravity is shifting beyond the model itself. Mistral’s record round shows Europe is intent on building a sovereign full stack. Meta is betting that personal agents will turn AI into a mainstream consumer product. And the Navier–Stokes controversy, alongside the Anthropic resignation, makes clear that as these systems take on more consequential work, the questions of trust, credit, and control are growing just as fast as the capabilities.