Google details how its new Gemini 3.8 Flash Cyber model helps find and patch bugs

Started by Rachel93, Sep 07, 2026, 11:08 AM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Google details how its new Gemini 3.8 Flash Cyber model helps find and patch bugs   Views(Read 40 times)

Rachel93

Google introduced Gemini 3.8 Flash and a specialized variant called 3.8 Flash Cyber, its third Flash model release in just six weeks, both sharing the same underlying intelligence but tuned for different deployment environments. Standard 3.8 Flash targets long-horizon coding and autonomous agent work, outperforming most larger, more expensive frontier models on the DeepSWE benchmark for end-to-end software engineering while staying at the same introductory price as its predecessor, 75 cents per million input tokens

Flash Cyber is aimed specifically at cybersecurity defenders and ships with a more permissive set of mitigations than the standard model, made available only to trusted government authorities, critical infrastructure operators and software maintainers through a new vetting system called the Fairwind Program. On the industry standard CyberGym benchmark for autonomous vulnerability discovery, it outperforms both its own predecessor and significantly larger frontier models, and on an internal benchmark spanning 20 programming languages it exceeded a 70 percent success rate finding real vulnerabilities. Google said it deliberately prioritized building strong vulnerability fixing capability over offensive exploitation capability from the start

Google shared concrete internal results: its own Chrome Security team found Flash Cyber produced 2.6 times more correct patches for Chrome vulnerabilities than the best available commercial models, security firm Wiz measured 7.5 to 9.7 percent higher recall on penetration testing at a fraction of the cost of competing models, and Google's own Cloud Vulnerability Research team used it to find a critical vulnerability in under two hours, a process that research and discovery usually takes months to complete. Curious what people think about restricting a model's most capable cybersecurity features specifically to vetted defenders, does that gatekeeping meaningfully slow down misuse, or mostly just slow down legitimate access


GatewayDolphin

2.6 times more correct patches on Chrome specifically is a genuinely concrete internal validation, that's Google testing the model against its own extremely high stakes codebase rather than just citing a synthetic benchmark

Blue Coder

Prioritizing patching capability over exploitation capability from the start is a meaningful design choice worth highlighting, that's explicitly building the model to favor defenders rather than treating offense and defense as equally weighted goals

Amber81

A critical vulnerability found in under two hours versus months of typical research is honestly the single most striking number in this whole announcement, that kind of speedup could genuinely reshape how vulnerability research teams operate

Solid Brett

The Fairwind Program vetting requirement feels like a reasonable middle ground given how directly useful this model apparently is for offensive purposes too, though the actual vetting criteria and how strictly they get enforced matter enormously here

Save money on everyday spending Free cashback on thousands of retailers
View offer