Three models in one week: Gemini 3.8 Flash rewrites the price rules

Three models in one week: Gemini 3.8 Flash rewrites the price rules

September 2026: Claude Fable 5.1, Gemini 3.8 Flash and GPT-6 Astra all shipped within days. Google's little Flash impresses with its intelligence-to-price ratio, but each model has its own sweet spot.

Summer ends and AI labs go back to shipping. Three frontier releases in three days: Claude Fable 5.1 from Anthropic on September 1st, Gemini 3.8 Flash from Google on the 2nd, GPT-6 Astra from OpenAI on the 3rd. Anyone who had set their out-of-office came back to a changed landscape.

But the interesting story isn’t the volume of releases. It’s the direction: the real battleground is no longer “who wins the benchmarks” but at equal quality, how much do you pay.

Gemini 3.8 Flash: the price surprise

Built on top of 3.7 Flash, Google’s new model is optimized for agents and coding. One-million-token context window, 64K output, knowledge cutoff March 2026. Benchmarks are solid: 95.3% on GPQA Diamond (frontier-level), but 19.1% on Terminal-Bench 4.0 - far behind the flagships.

And here’s the point: Flash doesn’t compete with flagships on the same tasks. It competes in the “good enough, at a tenth of the price” tier. It’s the model you choose when you have high call volume, lightweight agents, repetitive tasks, and you don’t need the 10$/million-token brain. For most enterprise use cases, it’s all you need.

Claude Fable 5.1: king of long-horizon coding

Anthropic made a smart move. Fable 5.1 isn’t just better than its predecessor: it costs less. Cache reads (when the model reuses already-processed context) dropped 75%, to $0.25 per million tokens. Result: -25% cost on typical workloads, up to -45% on heavy agentic tasks.

The numbers speak clearly: 55.8% on Terminal-Bench 4.0 (agentic coding), 65% on Humanity’s Last Exam with tools (above everyone, Astra included). It’s the model you choose when you have a complex coding problem that requires hours of autonomous work.

Two important notes for European businesses: Fable 5.1 supports zero data retention for eligible customers, and includes the watermarking required by the EU AI Act (signed by Anthropic in July 2026 alongside 190 other signatories). The detection API is in private preview for regulators and qualified organizations.

GPT-6 Astra: raw power, limited access

OpenAI’s flagship is the strongest model on computer use: 72.6% on OSWorld 2.0, up from 65.7% of its predecessor GPT-5.6 Sol, with a 47% cut in time per task. It saturates FrontierMath Tier 4 (97.6%) and ExploitBench (100%). Price: $10/$50 per million tokens, aligned with Fable 5.1.

But there are three caveats. First: the 99.9% on ARC-AGI-3 depends on an expensive stateful harness; stateless API calls score much lower. Second: on Humanity’s Last Exam with tools, Astra scores 57.2%, below Fable 5.1’s 65%. Not a clean sweep across the board. Third: access is still limited to OpenAI’s Trusted Access Program, with gradual rollout.

Cybersecurity capabilities crossed the “Critical” threshold of OpenAI’s Preparedness Framework: exploit creation is blocked by default, and the Daybreak program will manage access to more advanced features.

The practical choice for businesses

Not everything needs a flagship. Today’s choice is clear:

  • Gemini 3.8 Flash for volume, lightweight agents, and repetitive tasks where cost per token is decisive
  • Claude Fable 5.1 for complex coding and knowledge work requiring long-horizon reasoning
  • GPT-6 Astra for advanced computer use and complex automation, if you can manage the limited access

The real winner of September, for businesses, isn’t the model with the highest score. It’s the one that offers the right intelligence-to-price ratio for your use case. And sometimes, the little Flash is more than enough.

This is exactly why we built AIDeskPro the way we did: plurality of engines as a principle. The landscape changes every week, September 2026 proved it, and no business should commit to a single LLM. We select, test, and integrate the best models, serve them from European endpoints with zero data retention, and transparently label those that aren’t European. You pick the right engine for each task.