
Three Gemini 3.8 Flash Releases in Six Weeks: Benchmark Scores Near Opus 5, but Real-World Use Tells a Different Story
Google has released its third Flash model in just six weeks.
In the early hours of September 2 Beijing time, Gemini 3.8 Flash officially launched. Just three weeks had passed since the release of 3.7 Flash, and barely a month and a half since 3.6 Flash. Google appears determined to push its rapid-iteration strategy to the limit—but it also raises a question: at this pace, will the 3.x version numbers even last until Gemini 4 arrives?
Based on Google’s official benchmarks, Gemini 3.8 Flash has an impressive scorecard.
On DeepSWE v1.1, a long-horizon software engineering benchmark, 3.8 Flash scored 73.7%—just 0.3 percentage points behind Claude Opus 5 at 74.0%. Gemini 3.6 Flash, released just six weeks earlier, managed only 49%. On Terminal-Bench 2.1, 3.8 Flash slightly edged out Opus 5, scoring 89.4% versus 89.1%. On HLE-Verified, a broad reasoning benchmark, it scored 54.9%, again narrowly ahead of Opus 5 at 54.4%.
Even more surprising is the pricing. Gemini 3.8 Flash costs exactly the same as 3.7 Flash: $0.75 per million input tokens and $3.75 per million output tokens. With Opus 5 priced at $5 per million input tokens and $25 per million output tokens, Gemini 3.8 Flash comes in at roughly 15% of the price.
Google explained in its official blog that 3.8 Flash can handle harder tasks because it simply works harder. When faced with complex problems, it performs more reasoning steps, repeatedly calls tools, and checks its own work. The model took first place on eight of 14 benchmarks and scored 59 on the Artificial Analysis Intelligence Index—matching GPT-5.6 Sol and Grok 4.6 in their non-maximal reasoning configurations.
Good benchmarks don't necessarily mean a model is good to use.
In testing by iFanr, 3.8 Flash generated a 3D helicopter model in less than two minutes. But once opened, the result was noticeably rough: the body details, component relationships, and movement were all fairly crude. It could fly, but its resemblance to an H145 helicopter was about as close as a chicken-shaped pastry is to an actual chicken.
A separate 3D water-flow simulation took roughly the same two minutes. The result worked—but the volcano was visibly leaking water.
The subtlety here is that these outputs aren't outright failures. The engineering task gets completed, the code generally runs, and the required functions are there. But the gap between the result and Google's polished demos is visible to the naked eye.
Early user feedback shows a similar split: response speed is widely praised, but in real-world coding and generation tasks, the model can still be too quick to interpret requirements, producing something that looks complete on the surface while falling apart in the details.
The problem lies partly in how the demos are run.
In Google's official demonstrations, 3.8 Flash operates inside Antigravity through a continuous instruction loop, while also calling Nano Banana to generate textures. In other words, Google is showing the model's ceiling when supported by an entire technology stack. Users, meanwhile, encounter something much closer to its floor when the model works on its own.
Then there's the more practical question: the bill.
Artificial Analysis found that 3.8 Flash's total cost for completing the same task rose by 40% in testing—the model consumed 30% more output tokens per task. The price sheet hasn't changed. The bill has.
And the current discounted pricing only lasts through the end of this year. Starting January 1, 2027, the unit price will double. More usage first, higher prices later—the two increases together reveal the full pricing trajectory.
Repeated delays to Gemini 3.5 Pro have forced Google to temporarily step back from the Pro market. After taking over as head of DeepMind, Koray Kavukcuoglu emphasized the need to accelerate execution. Until Gemini 4 arrives, Flash has to hold the line across users, products, developers, and revenue.
The timing of 3.8 Flash's release is also telling.
Just one day earlier, Anthropic had completed the limited rollout of Fable 5.1 and Mythos 5.1. On the same day, Meta released Muse Spark 1.3. All three giants pushed out model updates within virtually the same window.
The battleground is shifting from price per token to cost per task.
Although 3.8 Flash's per-task cost increased by 40%, its roughly $0.58 cost per task still sits near the frontier of the intelligence-to-cost curve. For enterprises looking to deploy general-purpose AI agents at scale, that price remains highly attractive.
Google also launched Gemini 3.8 Flash Cyber, a model specifically designed for cybersecurity. It achieved 47.2% on the CWE-Bench vulnerability-fixing benchmark, roughly matching Fable 5. The model is currently available only to trusted security teams through the Fairwind program.
Stepping away from the race for the most powerful model may be the easier business.
But it also means that until Gemini 4 actually arrives, Flash has to carry the burden of representing Google's technical firepower.
That makes 3.8 Flash an unusually contradictory model. It needs to be fast and cheap enough to be embedded in Search, smartphones, and the daily lives of a billion people. At the same time, it needs to work for long stretches, repeatedly call tools, and behave more like a flagship model.
The ideal sounds great. Reality is messier.
Caught between two fundamentally opposing goals, 3.8 Flash can only play the role of a Pro model for so long—at least on the official benchmark charts.