Early this morning, Google made another move, with Gemini 3.8 Flash going live without warning.
It has been only three weeks since the last-generation 3.7 Flash was released - this is the third Flash series model that Google has launched in just six weeks, with a new generation being released every two weeks on average, a pace that was almost unimaginable in the large model circle in the past.
What's even more surprising is the price. The input is $0.75 per million tokens, and the output is $3.75, which is completely consistent with the 3.7 Flash. The "bargain" label is firmly stuck on its face. Also making a simultaneous debut is the Gemini 3.8 Flash Cyber, which specializes in network security and is open to trusted defenders through the brand-new Fairwind program.
The official data is impressive: DeepSWE's long-range programming has caught up with Claude Opus 5, with Terminal-Bench 2.1 (terminal operation benchmark) approaching the 90% mark, and its HLE inference score has even surpassed that of GPT-5.6 Sol. Yao Shunyu, a core member of Google DeepMind, left a thought-provoking comment: "For the model, this is just a small step; but for RSI (Recursive Self-Improvement), it's a giant leap."
Although, benchmarks are just benchmarks. When this "price-performance killer" pushed by Google fell into the hands of real developers, community feedback was far from as optimistic as the official press release - doing things quickly and delivering a lot is still not the same as truly doing them well.
Gemini 3.8 Flash outperforms top flagship phones at a bargain price
First, let's look at the hardware lineup from Google.
Gemini 3.8 Flash is positioned as the "smartest Flash model to date", with a focus on enhancing long-cycle software engineering, autonomous Agent tasks, and complex multi-step reasoning. The official benchmark test scores are impressive:
First, programming capabilities are approaching the industry ceiling. In the long-cycle software engineering test of DeepSWE v1.1, 3.8 Flash achieved 73.7%, while its predecessor 3.7 Flash scored 65.3%, directly matching Claude Opus 5's 74.0%. Independent verification by third-party agency Datacurve also showed that both had a pass rate of 74%, but 3.8 Flash had an average cost per question of only $2.36, while Opus 5 cost $11.84 - a difference of five times.

On Terminal-Bench 2.1, 3.8 Flash achieved 89.4%, surpassing Opus 5's 89.1%, becoming the first Flash-level model to approach the 90% threshold on this benchmark.

Second, cross-disciplinary reasoning has entered the top tier. In the HLE-Verified (Human-like Exam) test, 3.8 Flash scored 54.9%, slightly higher than Opus 5's 54.4% and GPT-5.6 Sol's 54.5%. With its multi-step reasoning capabilities covering STEM, humanities, and professional fields, a model at the price of Flash can now sit at the same table as flagship models.

Third, vertical professional Agent scenarios surpassed the flagship model. In the Vals financial Agent V2 evaluation, 3.8 Flash achieved 61.4%, exceeding 3.7 Flash's 59.0%, Opus 5's 58.6%, and GPT-5.6 Sol's 53.8%.

Based on the Harvey legal agent benchmark, it achieved a completion task rate of 10.0%, while Opus 5 had 6.7%.

Behind these achievements lies a core design choice: 3.8 Flash is "more diligent" when tackling complex tasks. When faced with difficult problems, it no longer rushes to provide answers, but instead takes more reasoning steps, repeatedly invokes tools, and verifies its results. In the DeepSWE task, 3.7 Flash averaged 107,000 tokens and 125 steps, while 3.8 Flash averaged 143,000 tokens and 166 steps. The cost is that while the unit price remains unchanged, the actual cost of complex tasks has risen from $0.40 to approximately $0.58.
As for the Cyber version, it exceeded GPT-5.5-Cyber with 86.2% Pass@1 on the CyberGym vulnerability discovery benchmark. In the CWE-Bench automatic repair test, it achieved a score of 47.2%, nearly tying with the industry-leading 47.8%, but at a significantly lower cost. In practical applications, the Chrome security team used it to generate a number of correct patches that was 2.6 times that of other commercial models. The Google Cloud team also used it to locate a critical vulnerability in under 2 hours, which would have taken several months to discover previously. However, this model is not publicly released and is only available to trusted defenders through the Fairwind program.
Benchmarks are one thing, but real-world testing by the community is another story altogether. When this so-called "value-for-money killer" fell into the hands of actual developers, the feedback was far from as optimistic as the official press release.
Top-Scoring Student Stumbles on Exam Day?
In an official Google demo, 3.8 Flash can be used to create a 3D magic castle and a DOS-style Google Maps interface - the effect is cool and realistic. However, when it falls into the hands of real developers, the style changes immediately.
Firstly, the configuration used for official demonstrations is impossible for ordinary users to obtain.
Ai Fan Er let 3.8 Flash use Three.js to build an Airbus H145 helicopter, and in 109 seconds, something was indeed produced - but upon closer inspection, the relationships between the aircraft's components were rough, and the movement was awkward; the 3D water flow simulation of a volcano was even "leaking". With the same prompt, the official version was able to produce a castle, while the user could only get a semi-finished product.
The difference lies in the fact that the official example is not just about throwing in the prompt and being done with it, but rather the model continuously loops through instructions within the Antigravity platform, while also invoking Nano Banana to generate textures. In other words, the 73.7% score is the result of the entire system, which includes the "model + Agent framework + multi-round verification + tool invocation", rather than the model acting alone.
Secondly, Flash's boast of being "fast" also cannot withstand scrutiny.
Developer Aditya used 3.8 Flash and Kimi K3 to develop the same Sticky Ball game - while Flash was indeed faster and more token-efficient, the game's motion mechanism and movement performance were ultimately inferior.

AI Pulse Daily's New York City market scene test was even more of a "public execution": with the same prompt, 3.8 Flash only generated 6 trees - fewer than the 10 generated by 3.7 Flash three weeks ago - while Opus 5 managed to pack in 1840. The prompt explicitly requested pedestrian crossing stripes and underwater bokeh, but these were omitted, and instead, it added 5 camera presets and 3 post-processing switches - failing to deliver what users wanted while adding unnecessary features.

Furthermore, they are impressive in short sprints but struggle with long-distance endurance.
On Terminal-Bench 4.0, the 3.8 Flash scored 19.1%, while the Opus 5 scored 51.8%; in the OSWorld 2.0 computer operation task, the Flash scored 59.0%, and the Opus 5 scored 75.4%. In the Artificial Analysis comprehensive intelligence rankings, the 3.8 Flash ranked eighth with a score of 59, in the same tier as the Kimi K3 and GLM-5.3 - while it can almost reach second place in single-item coding tasks, its lack of foundation is fully exposed when faced with long tasks that require endurance, multiple steps, and cross-domain capabilities.
This raises a more pressing question: given the significant disparity in actual performance, why is Google in such a hurry to release 3.8 Flash?
Google Releases Three Flash Devices in Six Weeks, But Why the Rush?
In six weeks, three Flash models have been released, which on the surface appears to be a result of rapid technological iteration, but in reality is a sign of desperation.
The flagship Pro has been delayed, with its release pushed back repeatedly. Pichai revealed in May that the new Pro might arrive "in about a month", but Google's internal Gemini 3.5 Pro plan was ultimately scrapped - the main reason being that its improvements over the Flash series were not significant enough. The delay from May to September has been so long that people have stopped asking about it. The more distant Gemini 4 has completed part of its pre-training, but is still stuck in the post-training phase, and its official release is nowhere in sight.
Core technical talent is bleeding away amid organizational turmoil. Over the past month and a half, Google has experienced a wave of personnel changes. Key technical figures including Noam Shazeer and Jeff Dean departed in succession this summer—even Jeff Dean, the "binary capability holder," left. Google subsequently reshuffled DeepMind's management, with Koray Kavukcuoglu taking on broader managerial responsibilities. The new leadership's directive is unambiguous: further accelerate the pace of model development and release.
The scale of Flash is relatively small, and the computing power required for modification and retraining is much lower than that of the Pro series. According to a disclosure by the Wall Street Journal, within Google, multiple teams can simultaneously try out different fine-tuning schemes on Flash, completing a round of iteration in about three weeks. In contrast, using the Pro series involves more GPUs, longer training times, and cross-team resource coordination for each adjustment, with trial and error costs being vastly different.
In other words, Google's ability to release three Flash models in just six weeks is not due to a sudden technological breakthrough, but rather because the Flash line itself is more suited to a "small step, quick pace" approach. With the flagship models facing production difficulties, the Flash naturally became the only product line that could continuously release new versions.
So, under the dual pressure of struggling with its flagship and Flash being naturally suited for "small steps, quick runs", Google chose to "achieve big things with small steps" and let Flash take the lead. With three products in six weeks, high-frequency iterations, it can move as fast as it wants.
Not Focusing on Being the Best May Actually Lead to Greater Gains
Benchmarks that surpass expectations and actual test results that are disappointing - when viewed together, these contradictions actually provide a clearer picture of Flash's true situation.
It's not yet a flagship, but it's already taking on flagship tasks; it's not yet all-powerful, but it's cheap enough, fast enough, and can be iterated upon continuously. This shows that Flash still has a gap to bridge to become a truly flagship-level all-around model, but for Google now, it's sufficient.
While Pro is long overdue and Gemini 4 is still on the way, Flash is currently Google's fastest card to play. With three updates in six weeks, it appears that Flash is surging forward, but in reality, Google is racing against time.
Whether it can win the final game still depends on Pro and Gemini 4, but at least for now, Flash is holding the table for Google.
