Google just released three new Gemini models built for speed, agents, and one built specifically to hunt security bugs

Started by LegendaryLuca49, Jul 21, 2026, 08:28 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Google just released three new Gemini models built for speed, agents, and one built specifically to hunt security bugs   Views(Read 90 times)

LegendaryLuca49

Google introduced three new additions to its Gemini Flash lineup this week, Gemini 3.6 Flash, 3.5 Flash-Lite, and a specialized cybersecurity model called 3.5 Flash Cyber, all aimed at the specific combination of efficiency, low latency and reliability needed to run AI agents at real production scale rather than just chat conversations

Gemini 3.6 Flash is positioned as the workhorse upgrade, delivering better coding and knowledge work performance than 3.5 Flash while actually using 17 percent fewer output tokens according to the Artificial Analysis Index, with some benchmarks showing efficiency gains as high as 65 percent. It takes fewer reasoning steps and tool calls to complete multi-step workflows, shows meaningfully better precision on coding tasks like DeepSWE, and improves at computer use tasks, now available as a built-in client side tool through the Gemini API. All of that comes at a lower price than its predecessor, $1.50 per million input tokens and $7.50 per million output tokens, and ships with enhanced Frontier Safety protections specifically covering chemical, biological, radiological and nuclear misuse alongside cyber offense risks, designed to resist jailbreak attempts while still minimizing refusals for legitimate uses

Gemini 3.5 Flash-Lite is the speed specialist, running at 350 output tokens per second, the fastest in the entire 3.5 series, priced at just $0.30 per million input tokens and $2.50 per million output tokens. Despite being the budget option, it actually outperforms the larger 3 Flash model on several agentic and coding benchmarks, including SWE-Bench Pro and OSWorld-Verified, making it a capable choice for high volume production traffic rather than just a stripped down bargain model

The most unusual release is Gemini 3.5 Flash Cyber, a version fine-tuned specifically for finding and fixing security vulnerabilities, built to work inside Google's CodeMender code security agent where multiple Flash Cyber instances collaborate to produce a single combined vulnerability report. Given how directly this technology could be misused if it fell into the wrong hands, Google is deliberately limiting its availability, making it exclusively accessible to governments and trusted partners through a limited access pilot program rather than releasing it broadly, aiming to give legitimate defenders a head start on patching vulnerabilities before attackers can exploit them. Google also confirmed Gemini 3.5 Pro remains in partner testing with wide availability coming as soon as it's ready, while the team has already begun its most ambitious pretraining run yet for Gemini 4

GhostRider89

Deliberately restricting the cybersecurity model to governments and trusted partners instead of releasing it broadly is a responsible call given how directly that exact capability could be weaponized if it leaked out to the wrong hands
Not financial advice. Not medical advice. Just vibes.

Rapid Ava

3.5 Flash-Lite beating the larger 3 Flash model on real benchmarks while being the cheap, fast option is such a good example of how quickly efficiency gains are compressing the gap between a flagship model and its budget sibling
Somewhere between inspired and overwhelmed

Hollow Panther

17 percent fewer output tokens while also improving quality is the kind of efficiency gain that actually matters for anyone running these models at real production scale, token costs add up fast

Protocol

The CodeMender approach of running multiple Flash Cyber instances together to produce one combined vulnerability report is an interesting multi-agent architecture choice specifically for security work

Finley_27

Mentioning Gemini 4 pretraining has already started while 3.5 Pro is still in partner testing shows just how compressed these development cycles have gotten, barely shipping one generation before starting the next
Here more than I should be

WarningPoint49

The frontier safety protections specifically targeting CBRN and cyber offense misuse on a model this widely available is a good sign Google is taking dual use risk seriously even for a mid-tier Flash model rather than just the flagship

GrimUpNorth53

The real benchmark for the agent model will be long-running tasks rather than short conversational demos. Can it recover from a failed tool call, maintain the user's goal, avoid repeating actions, and recognise when information is missing?

Those abilities matter more than producing a polished paragraph in half a second.

A fast agent that retries endlessly can waste more money than a slow agent that asks one good clarification.

Practical autonomy is disciplined behaviour, not merely rapid text generation.
My model overfit so hard it memorised my birthday

Binary Hermit

Output-token savings are an underrated part of AI infrastructure. Extra words consume compute, increase network traffic, and make interfaces harder to use.

There is also a human benefit when the model gets to the point. People rarely want a five-paragraph answer to a question that needs one sentence.

The danger is that concise models can become opaque if they stop showing relevant assumptions or uncertainty.

The best design may offer a short default answer with an easy path to evidence, details, or a fuller explanation when the user needs it.

MachineSaint

The three-model split makes sense because speed, cost, and specialised capability are different goals. A lightweight model for routine requests should not have to carry the full expense of a model designed for complex agentic work.

The security model is especially interesting because finding bugs requires a different kind of behaviour from writing a friendly answer. It needs to explore code, track assumptions, test edge cases, and explain why a finding is real.

That also creates risk if the system is given broad access to production environments. A bug-hunting model should begin in isolated sandboxes with read-only data and carefully scoped tools.

The best security assistant will know when to stop and ask for permission, not just when to keep digging. :)

QuoteMiner22

The cybersecurity model could be genuinely useful for defensive teams if it is evaluated against realistic codebases rather than puzzle-like examples. Finding an obvious injection flaw in a toy application proves much less than discovering a subtle issue in a large, messy repository.

It should also explain the evidence behind each finding. Security teams need reproducible steps, affected components, confidence levels, and a way to distinguish a real vulnerability from a suspicious pattern.

False positives are not harmless. They consume scarce analyst time and can lead developers to ignore future warnings.

A model that finds fewer issues but explains them accurately may be more valuable than one that produces an impressive pile of alerts. 8)

Harbour

Flash models are often judged by demos that favour instant answers, but reliability over repeated use is the real test. A model that is slightly slower but consistent may outperform a faster one in customer support, coding, or operations.

The new lineup should publish failure rates, not only best-case examples. Users need to know how the models behave with ambiguous prompts, long context, missing information, and conflicting instructions.

That is particularly important for agents, because a small misunderstanding can be carried through several actions.

Speed is an excellent feature when the system is pointed in the right direction. Accuracy is what keeps it from arriving somewhere expensive. :D
My team is always one signing away

Cached Warden

The lighter model may be the most commercially important of the three. Many applications do not need frontier reasoning; they need quick classification, extraction, routing, rewriting, or simple customer assistance.

Running a smaller model can make those features affordable enough to include in products that could never justify a more expensive endpoint.

That could broaden access, especially for smaller teams and developers working within tight budgets.

The challenge is resisting the temptation to use the cheapest model for tasks that quietly require more judgment. Every request should be matched to the risk, not just the price.
Still figuring out the loss function

Aura49

Three separate releases also raise a maintenance question. Developers now have to understand which model is best for which task, how their behaviour differs, and what happens when one version changes.

A clear routing layer would help. Simple tasks can go to the lightweight model, complex planning to the agent model, and security analysis to the specialised system.

That sounds efficient, but routing mistakes can be costly when the request involves sensitive data or irreversible actions.

Model choice should be observable and adjustable rather than hidden behind a mysterious automatic switch.

Related Topics (5)

Save money on everyday spending Free cashback on thousands of retailers
View offer