
Google’s Third Flash Release in Six Weeks
On September 2, Google introduced two new models: Gemini 3.8 Flash, a general-purpose reasoning and coding model, and Gemini 3.8 Flash Cyber, a specialized cybersecurity variant. It’s the company’s third Flash-tier release in six weeks, following Gemini 3.7 Flash roughly three weeks earlier — a release cadence that signals Google is treating its mid-tier model line as the primary battleground rather than reserving frequent updates for flagship models.
Gemini 3.8 Flash: More Reasoning, Same Price
Gemini 3.8 Flash is positioned as Google’s strongest coding and reasoning model at Flash-tier pricing — $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026, rising to $1.50/$7.50 in January 2027. Google reports it often approaches costlier frontier models on long-horizon software engineering benchmarks (DeepSWE v1.1) and scores 54.9% on HLE-Verified, a difficult multi-domain reasoning benchmark.
The mechanism behind the gains is notable: rather than a larger model, Google attributes the jump to the model working harder at inference time — taking more reasoning steps and calling tools iteratively, which means real-world token consumption (and cost) can scale up on complex tasks even though the sticker price hasn’t moved. Developers optimizing for cost still have the option to stay on 3.7 Flash or dial down effort levels.
Google is also positioning the model as an agentic coding tool through Google Antigravity, its agent-first development environment, showcasing demos like a fully playable browser-based game and a DOS-style clone of Google Maps generated from single prompts.
Gemini 3.8 Flash Cyber: Defense-First, Gated Access
The more distinctive release is Flash Cyber, a cybersecurity-specialized variant available only through Google’s new Fairwind Program, aimed at vetted defenders — government authorities, critical infrastructure operators, and software maintainers. Google explicitly says it prioritized patching and vulnerability-fixing capability over offensive exploitation capability, a deliberate defense-over-offense design choice.
On CyberGym, an industry-standard vulnerability discovery benchmark, Google reports frontier-level performance, and on an internal benchmark spanning 20 programming languages, a vulnerability discovery success rate above 70%. On automated patching (CWE-Bench, run by Collinear), Flash Cyber lands near the top of the cost-performance frontier — comparable pass rates to a leading frontier model at meaningfully lower cost.
Google cites real internal deployments: the Chrome Security team reportedly saw 2.6x more correct vulnerability patches versus larger commercial models; Wiz reports higher recall on penetration-testing benchmarks at a fraction of the cost of competing frontier models; and Google’s Cloud Vulnerability Research team says it found a critical vulnerability in under two hours using the model, a process that typically takes months.
Safety Posture
Both models ship with restrictions against misuse in CBRN (chemical, biological, radiological, nuclear) and cyber-offense domains, per Google’s Frontier Safety Framework. Flash Cyber carries a more permissive set of cyber-specific mitigations than the general-purpose model, which is precisely why it’s restricted to the vetted-partner program rather than shipped broadly. Google also reports meaningful gains in prompt-injection robustness as measured by Gray Swan’s benchmark.
Why It Matters
Two competitive signals worth watching. First, the release velocity itself: three Flash updates in six weeks is a pace aimed squarely at compressing the window competitors have to respond, and it puts pressure on every other lab’s mid-tier pricing and capability claims. Second, the Fairwind Program is a template other labs are likely to study: rather than restricting a powerful dual-use capability outright, Google is shipping a defense-optimized version to a vetted circle while keeping general access models more conservative — a middle path between “don’t ship it” and “ship it to everyone.”
For teams evaluating model providers, the practical takeaway is that “same price, more reasoning” claims need to be tested against real task token consumption, not list price alone — since effort-scaling models can quietly increase spend on harder workloads even when the per-token rate is unchanged.
Source: Introducing Gemini 3.8 Flash and 3.8 Flash Cyber, Google, September 2, 2026.
