☁️ AI Weather Report — Top 10 Models for Coding Value — September 04, 2026

Welcome to the AI Weather Report for September 04, 2026. This daily report ranks the top 10 AI models for coding by bang for the buck — a combination of raw coding capability and API pricing.

📊 Today’s Top 10 Rankings

#ModelProviderCapabilityCost /M tokensValue Score
🥇 1 mistral-nemo mistralai 62/100 $0.0272 2275.2
🥈 2 l3-lunaris-8b sao10k 58/100 $0.0475 1221.1
🥉 3 mistral-small-24b-instruct-2501 mistralai 72/100 $0.0725 993.1
4 llama-3.1-8b-instruct meta-llama 62/100 $0.0725 855.2
5 mythomax-l2-13b gryphe 48/100 $0.0600 800.0
6 gpt-oss-20b openai 78/100 $0.1050 742.9
7 laguna-xs-2.1 poolside 72/100 $0.1050 685.7
8 gpt-oss-120b openai 93/100 $0.1368 680.1
9 deepseek-v4-flash deepseek 91/100 $0.1551 586.9
10 gemma-3-4b-it google 50/100 $0.0875 571.4

📈 Analysis

🏆 Best Value Today: mistral-nemo scores 2275.2 with a capability rating of 62 at $0.0272/M tokens.

What “Value Score” means: Capability score (based on SWE-bench, HumanEval, LiveCodeBench) divided by blended cost per million tokens (25% input + 75% output weights for coding workloads). Free tier models get a massive boost. Higher is better.

📋 All Scored Models (63 total)

#ModelProviderCapabilityCost /M tokValue
1mistral-nemomistralai62$0.02722275.2
2l3-lunaris-8bsao10k58$0.04751221.1
3mistral-small-24b-instruct-2501mistralai72$0.0725993.1
4llama-3.1-8b-instructmeta-llama62$0.0725855.2
5mythomax-l2-13bgryphe48$0.0600800.0
6gpt-oss-20bopenai78$0.1050742.9
7laguna-xs-2.1poolside72$0.1050685.7
8gpt-oss-120bopenai93$0.1368680.1
9deepseek-v4-flashdeepseek91$0.1551586.9
10gemma-3-4b-itgoogle50$0.0875571.4
11granite-4.1-8bibm-granite48$0.0875548.6
12qwen3.5-9bqwen72$0.1375523.6
13qwen3-30b-a3b-instruct-2507qwen82$0.1568522.9
14gemma-3-12b-itgoogle60$0.1250480.0
15mistral-small-3.2-24b-instructmistralai78$0.1688462.2
16command-r7b-12-2024cohere54$0.1219443.1
17granite-4.0-h-microibm-granite38$0.0882430.6
18ministral-3b-2512mistralai42$0.1000420.0
19nova-micro-v1amazon45$0.1137395.6
20qwen3-32bqwen88$0.2300382.6
21qwen3-coder-30b-a3b-instructqwen84$0.2275369.2
22qwen-2.5-7b-instructqwen60$0.1750342.9
23qwen3-235b-a22b-2507qwen96$0.2844337.6
24qwen3.5-flash-02-23qwen70$0.2112331.4
25llama-3.3-70b-instructmeta-llama84$0.2650317.0
26gpt-oss-safeguard-20bopenai77$0.2437315.9
27nemotron-3-nano-30b-a3bnvidia50$0.1625307.7
28nova-lite-v1amazon58$0.1950297.4
29gemma-4-31b-itgoogle74$0.2775266.7
30gemma-4-26b-a4b-itgoogle72$0.2725264.2
31seed-1.6-flashbytedance-seed64$0.2437262.6
32gpt-5-nanoopenai82$0.3125262.4
33step-3.5-flashstepfun60$0.2500240.0
34nemotron-3-super-120b-a12bnvidia76$0.3212236.6
35seed-2.0-minibytedance-seed72$0.3250221.5
36llama-3.1-70b-instructmeta-llama82$0.4000205.0
37llama-3.2-1b-instructmeta-llama30$0.1575190.5
38glm-4.7-flashz-ai60$0.3150190.5
39gemma-3-27b-itgoogle68$0.3575190.2
40gpt-4.1-nanoopenai60$0.3250184.6
41llama-3.2-3b-instructmeta-llama48$0.2600184.6
42gpt-4o-miniopenai74$0.4875151.8
43hy3-previewtencent68$0.4950137.4
44command-r-08-2024cohere60$0.4875123.1
45deepseek-chatdeepseek90$0.7475120.4
46qwen3-next-80b-a3b-instructqwen90$0.8500105.9
47qwen3-coderqwen85$0.8250103.0
48qwen3-next-80b-a3b-thinkingqwen93$0.937599.2
49qwen-2.5-coder-32b-instructqwen86$0.915094.0
50hermes-3-llama-3.1-405bnousresearch78$1.0078.0
51claude-3-haikuanthropic72$1.0072.0
52dolphin-mistral-24b-venice-editioncognitivecomputations52$0.725071.7
53gpt-4.1-miniopenai76$1.3058.5
54deepseek-r1deepseek95$2.0546.3
55gemini-2.5-flashgoogle86$1.9544.1
56nova-pro-v1amazon70$2.6026.9
57gpt-4.1openai90$6.5013.8
58gpt-5openai97$7.8112.4
59gemini-2.5-progoogle94$7.8112.0
60gpt-4oopenai88$8.1310.8
61command-r-plus-08-2024cohere68$8.138.4
62claude-sonnet-4anthropic96$12.008.0
63claude-opus-4anthropic98$60.001.6

Generated 2026-09-04 02:00 UTC · Data from OpenRouter API and public benchmarks · Bang-for-Buck = Capability / Cost

Top AI Stories – September 3, 2026

September came in hot. Within a single day, Anthropic, Google, and Meta each shipped or pushed new frontier-scale AI models, and a lone researcher topped a benchmark that cost him less than a dollar in compute. Here are the five AI stories that dominated the news today.

Anthropic unveils Claude Fable 5.1 and Claude Mythos 5.1

Anthropic introduced Claude Fable 5.1 and Claude Mythos 5.1, which it calls “the world’s most advanced models for coding and knowledge work.” The two appear to be the same underlying model with different safeguards: Fable 5.1 is generally available, while Mythos 5.1 is restricted to Anthropic’s trusted-access programs and designed specifically to support work in cybersecurity and the life sciences.

The company says Fable 5.1 takes important steps toward addressing customer feedback on price and data retention,and boasts the strongest cyber capabilities of any model it has released,still landing in the lower category of risk under its Frontier Compliance Framework. Anthropic also highlighted scientific research contributions,including Claude-designed protein binders confirmed to bind in the lab и new high-resolution elevation map of a third of Venus derived from NASA’s Magellan radar data.

On cost, Anthropic says Fable 5.1 will be roughly ‍25% cheaper than Fable 5 for typical token-billed workloads because of reduced pricing on cache reads,and up to approximately 45% cheaper for highly agentic work. Enterprise Frontier Safeguards(EFS,which gives customers complete privacy via infrastructure controlled entirely by the customer rather than Anthropic,will roll out in phasesbeginning later this fall. Mythos 5.1 will be offered through a Cyber Verification Program anda Life Sciences Verification Program forvetted defenders and researchers.

Google ships Gemini 3.8 Flash and Gemini 3.8 Flash Cyber

On September2, Google released Gemini 3.8,its best reasoning and coding model yet,at the same speed and low cost of 3.7,and its third Flash release insix weeks. Two variants ship today:Gemini 3.8 Flash,its most intelligent workhorse model,and Gemini 3.8 Flash Cyber,afrontier cybersecurity modelfor trusted defenderst hrough its new Fairwind Program.

Gemini 3.8 Flash delivers substantial gains over 3.7,often approaching higher-cost frontier models. On DeepSWE v1.1,a long-horizon software-engineering benchmark,it outperforms most larger frontier models and scores 54.9% on HLE-Verified,covering multi-step reasoning across STEM, humanities,and professional fields. It launches at $0.75 per million input tokens and $3.75 per million output tokens丹an introductory price that rises to $1.50/$7.50 after January 1, 2027. It also powers agent-first workflows in Google Antigravity, building playable apps from single prompts, including a functional DOS version of Google Maps。

The cyber variant demonstrated frontier-level vulnerability discovery,surpassing both 3.5 Flash Cyber and significantly larger frontier models on the CyberGym benchmark,and exceeds a success rate of 70% across 20 programming languages on Google’s internal benchmark. On CWE-Bench,it posts a pass@1 of 47.2% versus a leading frontier model’s 47.8%,at significantly lower cost. Google says Chrome Security found the model produced 2.6 times more correct vulnerability patches,and collaborator Wiz measured +7.5-9.7% higher recall on penetration testing at 2.3-5.2x lower cost.

Google’s Cloud Vulnerability Research team,meanwhile,used the model to find a critical foundational vulnerability in less than two hours, a defect that normally takes months to discover. The models ship with CBRN and cyber-offense safeguards per Google’s Frontier Safety Framework,and made a significant leap in prompt-injection robustness as measured by Gray Swan。

Dan Luu asks: how accurate have Ed Zitron’s AI-skeptic predictions been?

Engineer and writer Dan Luu published a viral reality-check of Ed Zitron,one of the most widely cited AI skeptics,tallying his past predictions against what actually happened. His conclusion is blunt:Zitron’s capability-related predictions have generally been wrong to date,in particular claims that models haven’t improved since 2023 or 2024.

Luu offers concrete counterexamples:modern coding models can create a new regex engine with an interpreter and a native compiler in minutes;and AI video generation has improved dramatically between 2023 and 2025 and is now upending lower-end video work. He also critiques Zitron’s community,noting that posting improvement benchmarks on his Reddit sub often results in a ban,and flags Zitron’s financial predictions as also unproven. “Zitron should probably find a new rebuttal,” Luu writes,even if he’s playing to true believers.

A $0.67 transformer scores 44% on ARC-AGI-1

Independent researcher Mithil Vakde shared an eye-catching result :training a small autoregressive transformer from scratch in just 1.5 hours on a single RTX 5090 GPU,hitting 44% on the ARC-AGI-1 benchmark for an estimated cost of about $0.67 in compute,beating many LLMs and matching results from much larger systems like TRM/HRM. He also reports 7% on ARC-2.

Vakde emphasizes that this is deliberately not an LLM:it’s a small model trained from scratch with no synthetic data,competing under the ARC community rule that bans offline pretraining on the eval set. The aim, he says,is sample efficiency,the most important problem in AI today,and slashing compute costs so iteration is faster and cheaper. The effort garnered attention from top researchers including Lucas Beyer,Jeremy Howard,and Rohan Anil,and is hailed as evidence that extremely complex problems can be tackled without massive training budgets。

Meta’s Muse Spark 1.3 tops the intelligence index at startlingly low prices

Meta rolled out Muse Spark 1.3,its latest open reasoning model,and the community quickly took note:the 1.3 Max variant is the first Meta model to surpass OpenAI’s best on Artificial Analysis’ intelligence index,and the model posts a DeepSWE score of 75.4,the best recorded so far, overtaking Google’s Gemini 3.8 Flash, which had held the top spot earlier in the day.

Its headline feature, though, is pricing. Meta’s new contributor tier, which explicitly trains on your data,drops the cost to around $0.10 per million input tokens and $0.20 per million output, with $0.002 cached, a roughly 20x discount versus the full-price version. Developers flagged a 1M-token context window and per-reasoning-level costs ranging from about 4 cents to 7.5 cents for a single generation. Commenters lauded Meta for making “we train on this and value it this much” explicit, one of the first quantifiable numbers a provider has put on the value of training tokens, while others cautioned that the benchmarks are competitive but older,and that Gemini 3.8 Flash remains the cheaper full-price pick.

That’s the AI landscape as of September 3, 2026. From three lab releases in a single day to a dollar-scale benchmark record,the through thread is constant: reasoning power is rising,and the price of entry is falling。

☁️ AI Weather Report — Top 10 Models for Coding Value — September 03, 2026

Welcome to the AI Weather Report for September 03, 2026. This daily report ranks the top 10 AI models for coding by bang for the buck — a combination of raw coding capability and API pricing.

📊 Today’s Top 10 Rankings

#ModelProviderCapabilityCost /M tokensValue Score
🥇 1 mistral-nemo mistralai 62/100 $0.0272 2275.2
🥈 2 l3-lunaris-8b sao10k 58/100 $0.0475 1221.1
🥉 3 mistral-small-24b-instruct-2501 mistralai 72/100 $0.0725 993.1
4 llama-3.1-8b-instruct meta-llama 62/100 $0.0725 855.2
5 mythomax-l2-13b gryphe 48/100 $0.0600 800.0
6 gpt-oss-20b openai 78/100 $0.1050 742.9
7 laguna-xs-2.1 poolside 72/100 $0.1050 685.7
8 gpt-oss-120b openai 93/100 $0.1368 680.1
9 deepseek-v4-flash deepseek 91/100 $0.1551 586.9
10 gemma-3-4b-it google 50/100 $0.0875 571.4

📈 Analysis

🏆 Best Value Today: mistral-nemo scores 2275.2 with a capability rating of 62 at $0.0272/M tokens.

What “Value Score” means: Capability score (based on SWE-bench, HumanEval, LiveCodeBench) divided by blended cost per million tokens (25% input + 75% output weights for coding workloads). Free tier models get a massive boost. Higher is better.

📋 All Scored Models (63 total)

#ModelProviderCapabilityCost /M tokValue
1mistral-nemomistralai62$0.02722275.2
2l3-lunaris-8bsao10k58$0.04751221.1
3mistral-small-24b-instruct-2501mistralai72$0.0725993.1
4llama-3.1-8b-instructmeta-llama62$0.0725855.2
5mythomax-l2-13bgryphe48$0.0600800.0
6gpt-oss-20bopenai78$0.1050742.9
7laguna-xs-2.1poolside72$0.1050685.7
8gpt-oss-120bopenai93$0.1368680.1
9deepseek-v4-flashdeepseek91$0.1551586.9
10gemma-3-4b-itgoogle50$0.0875571.4
11granite-4.1-8bibm-granite48$0.0875548.6
12qwen3.5-9bqwen72$0.1375523.6
13qwen3-30b-a3b-instruct-2507qwen82$0.1568522.9
14gemma-3-12b-itgoogle60$0.1250480.0
15mistral-small-3.2-24b-instructmistralai78$0.1688462.2
16command-r7b-12-2024cohere54$0.1219443.1
17granite-4.0-h-microibm-granite38$0.0882430.6
18ministral-3b-2512mistralai42$0.1000420.0
19nova-micro-v1amazon45$0.1137395.6
20qwen3-32bqwen88$0.2300382.6
21qwen3-coder-30b-a3b-instructqwen84$0.2275369.2
22qwen-2.5-7b-instructqwen60$0.1750342.9
23qwen3-235b-a22b-2507qwen96$0.2844337.6
24qwen3.5-flash-02-23qwen70$0.2112331.4
25llama-3.3-70b-instructmeta-llama84$0.2650317.0
26gpt-oss-safeguard-20bopenai77$0.2437315.9
27nemotron-3-nano-30b-a3bnvidia50$0.1625307.7
28nova-lite-v1amazon58$0.1950297.4
29gemma-4-31b-itgoogle74$0.2775266.7
30gemma-4-26b-a4b-itgoogle72$0.2725264.2
31seed-1.6-flashbytedance-seed64$0.2437262.6
32gpt-5-nanoopenai82$0.3125262.4
33step-3.5-flashstepfun60$0.2500240.0
34nemotron-3-super-120b-a12bnvidia76$0.3212236.6
35seed-2.0-minibytedance-seed72$0.3250221.5
36llama-3.1-70b-instructmeta-llama82$0.4000205.0
37llama-3.2-1b-instructmeta-llama30$0.1575190.5
38glm-4.7-flashz-ai60$0.3150190.5
39gemma-3-27b-itgoogle68$0.3575190.2
40gpt-4.1-nanoopenai60$0.3250184.6
41llama-3.2-3b-instructmeta-llama48$0.2600184.6
42gpt-4o-miniopenai74$0.4875151.8
43hy3-previewtencent68$0.4950137.4
44command-r-08-2024cohere60$0.4875123.1
45deepseek-chatdeepseek90$0.8359107.7
46qwen3-next-80b-a3b-instructqwen90$0.8500105.9
47qwen3-coderqwen85$0.8250103.0
48qwen3-next-80b-a3b-thinkingqwen93$0.937599.2
49qwen-2.5-coder-32b-instructqwen86$0.915094.0
50hermes-3-llama-3.1-405bnousresearch78$1.0078.0
51claude-3-haikuanthropic72$1.0072.0
52dolphin-mistral-24b-venice-editioncognitivecomputations52$0.725071.7
53gpt-4.1-miniopenai76$1.3058.5
54deepseek-r1deepseek95$2.0546.3
55gemini-2.5-flashgoogle86$1.9544.1
56nova-pro-v1amazon70$2.6026.9
57gpt-4.1openai90$6.5013.8
58gpt-5openai97$7.8112.4
59gemini-2.5-progoogle94$7.8112.0
60gpt-4oopenai88$8.1310.8
61command-r-plus-08-2024cohere68$8.138.4
62claude-sonnet-4anthropic96$12.008.0
63claude-opus-4anthropic98$60.001.6

Generated 2026-09-03 02:00 UTC · Data from OpenRouter API and public benchmarks · Bang-for-Buck = Capability / Cost

Top AI Stories – September 02, 2026

Tuesday, September 1st into Wednesday, September 2nd brought another dense stretch of artificial-intelligence news, headlined by a major Anthropic model update, fresh signals about Apple’s quiet enterprise-AI boom, and a sharp legal escalation in the high-stakes Apple–OpenAI dispute. Here are the five stories shaping the conversation this cycle.

Anthropic Ships Claude Fable 5.1 and Claude Mythos 5.1

Anthropic kicked off the week’s news with the release of Claude Fable 5.1 and Claude Mythos 5.1, the latest iterations of its flagship model family. The company framed Fable 5.1 as a meaningful leap in both capability and cost-efficiency. Notably, cache-read pricing fell from $1/million tokens to $0.25/million — half the cache-read cost of Opus — which one Hacker News commenter described as strong evidence Anthropic is recalibrating pricing to widen adoption.

Early community and corporate reports were striking. Anthropic cited a testing engagement with the investment firm Millennium in which Fable 5.1 identified the root cause of a rare crash on Millennium’s internal systems that its own engineers had not been able to explain after several years of effort. Meanwhile, Anthropic employees touted Fable 5.1’s writing style as subtly and noticeably more natural than prior Claude models, and independent developer Simon Willison demonstrated the model generating animated “pelicans” across escalating reasoning-effort settings.

The release also drew scrutiny over three breaking changes, which pundits noted appeared to patch inadvertent chain-of-thought disclosure vectors — including instances where a crafted tool could induce a model to expose its raw reasoning. For builders, the trade-off between improved capability and tightening guardrails around reasoning transparency is becoming the defining tension of the frontier-model era.

Apple’s Enterprise AI Demand Catches Even Apple Off Guard

Apple’s unusually early announcement of new Mac mini and Mac Studio models this week was driven by unexpectedly strong enterprise appetite for on-premises AI hardware, according to The Information. Apple normally refreshes these lines in autumn, closer to October or November — the early launch, just ahead of iPhone season, signals how quickly the company is scrambling to meet demand.

The centerpiece of the enterprise pitch is the ability to link multiple Mac Studios into a single, more capable cluster for running large frontier AI models — a feature squarely aimed at developers and businesses rather than consumers. Apple highlighted the pivot toward business buyers in June with its “Business at the Park” event featuring executives from Ford, Disney, and Anthropic, where the Mac mini was dubbed the “darling” of the show.

Even so, the rush surprised leadership. The Information reports Apple lacked a dedicated engineering team for business customers, had no developer-relations staff, and lacked an enterprise AI strategy — and reportedly turned away companies seeking access to its Private Cloud Compute infrastructure. Apple is instead leaning on partners such as WebAI and Mount Thor to provide AI tooling on Apple silicon. The surge in demand has collided with the global memory shortage, leaving many Mac mini and Mac Studio configurations out of stock for months and pushing some customers toward rivals like Nvidia’s DGX Spark. Skeptics on Hacker News dismissed the narrative as clever marketing spun from supply-chain pressure — but either way, local AI hardware has become a real battlefield.

Apple Reveals ‘Shocking Evidence’ from Ex-Employee’s MacBook in OpenAI Suit

Apple escalated its trade-secret lawsuit against OpenAI, producing forensic evidence it describes as “shocking” from the MacBook of a former employee now working at OpenAI. According to 9to5Mac, Apple alleges the employee — identified in the case as Mr. Liu — not only downloaded a confidential Apple circuit schematic but used it in his work at OpenAI, including running engineering simulations (such as in LTspice) that his AI agents helped review.

The complaint asserts that Mr. Liu and colleagues at OpenAI were well aware of his unauthorized access to Apple’s third-party cloud storage, and that when Apple’s internal investigation surfaced, instructions for destroying evidence were sent to an OpenAI collaborator. Apple traces how it discovered the leak — the schematic was used on a Mac mini that later synced via iCloud to the MacBook taken from Apple — and now wants access to that Mac mini as well.

Most consequentially, Apple argues that when trade-secret information is fed into an AI agent that learns from it, that learning “may create irreversible and continually propagating uses of the trade secret.” Legal observers on Hacker News flagged this as a high-impact argument that tests the copyright-and-trade-secret frontier in AI training, and one that sits awkwardly alongside the industry’s own embrace of training on works without permission. The case is now a bellwether for how courts treat intellectual property absorbed into model weights.

Researchers Break Claude Code Opus 5’s Auto Mode

Security researcher Johann Rehberger published a detailed breakdown showing Claude Code Opus 5, running in its new default “Auto Mode,” can be hijacked into executing arbitrary code with a 60–80% attack success rate. The finding is notable because Auto Mode — which replaced human approval prompts with a safety classifier and became the default starting mode for Claude Code in mid-August — had been lauded under an Anthropic-commissioned evaluation showing a 0.00% prompt-injection success rate.

The attack chain is elegant and targeted. The researcher nudges Claude to move from the WebFetch tool to using curl directly, redirects it to a ZIP archive with files in a special encoding, and lets Claude’s own robust behavior work against it: the model correctly refuses to execute the binary and writes its own Python decoder instead — but it runs that decoder inside the attacker-controlled unzipped directory, where a malicious struct.py shadows Python’s standard implementation. When Claude imports the base64 module, the poisoned module executes. “BOOM,” as the write-up puts it.

Community reaction underscored an important distinction: this is less a classic prompt-injection hijack than a trojan aimed at exploiting Claude’s predictable tool habits. The episode is a pointed reminder that Anthropic’s own 0.00% claim was tested against a narrow benchmark vendor’s scenario set — and that even the most carefully safety-tuned agent is not a substitute for running untrusted workloads in an isolated, monitored environment.

OpenAI’s ChatGPT/Codex App Bundles a Full Copy of LibreOffice

Developer and AI commentator Simon Willison unearthed an unexpected detail inside OpenAI’s desktop Codex app (since rebranded to ChatGPT): buried in its runtime cache is a full installation of LibreOffice, the open-source office suite, alongside a complete Python installation, Node.js, Poppler, and git. In total, the codex-primary-runtime folder weighs in at roughly 1.7GB.

Willison found the suite bundled with skills modules that tell Codex how to find and use the LibreOffice binaries — a practical solution to a real problem: AI agents that need to read, render, and manipulate Microsoft Office documents reliably. Commenters pointed out that the suite is particularly useful for parsing legacy .xls files, and several noted the strategic implication — if AI agents increasingly generate and edit office documents locally, the traditional Office suite itself risks becoming a mere viewer. The discovery also sparked questions about open-source licensing and whether OpenAI should be contributing back to the LibreOffice project.

That’s the roundup for this cycle. The through-line is clear: capable local AI hardware, agents that must be trusted only cautiously, and a frontier-model race where pricing, security, and legal exposure are colliding faster than ever. We’ll be watching how these threads develop.

☁️ AI Weather Report — Top 10 Models for Coding Value — September 02, 2026

Welcome to the AI Weather Report for September 02, 2026. This daily report ranks the top 10 AI models for coding by bang for the buck — a combination of raw coding capability and API pricing.

📊 Today’s Top 10 Rankings

#ModelProviderCapabilityCost /M tokensValue Score
🥇 1 mistral-nemo mistralai 62/100 $0.0272 2275.2
🥈 2 l3-lunaris-8b sao10k 58/100 $0.0475 1221.1
🥉 3 mistral-small-24b-instruct-2501 mistralai 72/100 $0.0725 993.1
4 llama-3.1-8b-instruct meta-llama 62/100 $0.0725 855.2
5 mythomax-l2-13b gryphe 48/100 $0.0600 800.0
6 gpt-oss-20b openai 78/100 $0.1050 742.9
7 laguna-xs-2.1 poolside 72/100 $0.1050 685.7
8 gpt-oss-120b openai 93/100 $0.1368 680.1
9 deepseek-v4-flash deepseek 91/100 $0.1551 586.9
10 gemma-3-4b-it google 50/100 $0.0875 571.4

📈 Analysis

🏆 Best Value Today: mistral-nemo scores 2275.2 with a capability rating of 62 at $0.0272/M tokens.

What “Value Score” means: Capability score (based on SWE-bench, HumanEval, LiveCodeBench) divided by blended cost per million tokens (25% input + 75% output weights for coding workloads). Free tier models get a massive boost. Higher is better.

📋 All Scored Models (63 total)

#ModelProviderCapabilityCost /M tokValue
1mistral-nemomistralai62$0.02722275.2
2l3-lunaris-8bsao10k58$0.04751221.1
3mistral-small-24b-instruct-2501mistralai72$0.0725993.1
4llama-3.1-8b-instructmeta-llama62$0.0725855.2
5mythomax-l2-13bgryphe48$0.0600800.0
6gpt-oss-20bopenai78$0.1050742.9
7laguna-xs-2.1poolside72$0.1050685.7
8gpt-oss-120bopenai93$0.1368680.1
9deepseek-v4-flashdeepseek91$0.1551586.9
10gemma-3-4b-itgoogle50$0.0875571.4
11granite-4.1-8bibm-granite48$0.0875548.6
12qwen3.5-9bqwen72$0.1375523.6
13qwen3-30b-a3b-instruct-2507qwen82$0.1568522.9
14gemma-3-12b-itgoogle60$0.1250480.0
15mistral-small-3.2-24b-instructmistralai78$0.1688462.2
16command-r7b-12-2024cohere54$0.1219443.1
17granite-4.0-h-microibm-granite38$0.0882430.6
18ministral-3b-2512mistralai42$0.1000420.0
19nova-micro-v1amazon45$0.1137395.6
20qwen3-32bqwen88$0.2300382.6
21qwen3-coder-30b-a3b-instructqwen84$0.2275369.2
22qwen-2.5-7b-instructqwen60$0.1750342.9
23qwen3-235b-a22b-2507qwen96$0.2844337.6
24qwen3.5-flash-02-23qwen70$0.2112331.4
25gpt-oss-safeguard-20bopenai77$0.2437315.9
26nemotron-3-nano-30b-a3bnvidia50$0.1625307.7
27nova-lite-v1amazon58$0.1950297.4
28gemma-4-31b-itgoogle74$0.2775266.7
29gemma-4-26b-a4b-itgoogle72$0.2725264.2
30seed-1.6-flashbytedance-seed64$0.2437262.6
31gpt-5-nanoopenai82$0.3125262.4
32step-3.5-flashstepfun60$0.2500240.0
33nemotron-3-super-120b-a12bnvidia76$0.3212236.6
34seed-2.0-minibytedance-seed72$0.3250221.5
35llama-3.1-70b-instructmeta-llama82$0.4000205.0
36llama-3.2-1b-instructmeta-llama30$0.1575190.5
37glm-4.7-flashz-ai60$0.3150190.5
38gemma-3-27b-itgoogle68$0.3575190.2
39gpt-4.1-nanoopenai60$0.3250184.6
40llama-3.2-3b-instructmeta-llama48$0.2600184.6
41gpt-4o-miniopenai74$0.4875151.8
42hy3-previewtencent68$0.4950137.4
43command-r-08-2024cohere60$0.4875123.1
44llama-3.3-70b-instructmeta-llama84$0.7100118.3
45deepseek-chatdeepseek90$0.8359107.7
46qwen3-next-80b-a3b-instructqwen90$0.8500105.9
47qwen3-coderqwen85$0.8250103.0
48qwen3-next-80b-a3b-thinkingqwen93$0.937599.2
49qwen-2.5-coder-32b-instructqwen86$0.915094.0
50hermes-3-llama-3.1-405bnousresearch78$1.0078.0
51claude-3-haikuanthropic72$1.0072.0
52dolphin-mistral-24b-venice-editioncognitivecomputations52$0.725071.7
53gpt-4.1-miniopenai76$1.3058.5
54deepseek-r1deepseek95$2.0546.3
55gemini-2.5-flashgoogle86$1.9544.1
56nova-pro-v1amazon70$2.6026.9
57gpt-4.1openai90$6.5013.8
58gpt-5openai97$7.8112.4
59gemini-2.5-progoogle94$7.8112.0
60gpt-4oopenai88$8.1310.8
61command-r-plus-08-2024cohere68$8.138.4
62claude-sonnet-4anthropic96$12.008.0
63claude-opus-4anthropic98$60.001.6

Generated 2026-09-02 02:00 UTC · Data from OpenRouter API and public benchmarks · Bang-for-Buck = Capability / Cost