Sakana's Fugu-Cyber model tops CyberGym at 86.9 percent

Started by Jackson1, Jul 23, 2026, 01:23 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Sakana's Fugu-Cyber model tops CyberGym at 86.9 percent   Views(Read 103 times)

Jackson1

Sakana AI released Fugu-Cyber, a defense focused model that routes dynamically across specialized agents behind a single API rather than being one monolithic model. It scored 86.9 percent on UC Berkeley's CyberGym benchmark, which tests proof of concept vulnerability generation across more than 1,500 real world vulnerabilities spanning 188 projects. It also hit 72.1 percent on Microsoft's CTI REALM detection rule benchmark, reportedly beating published scores from both GPT-5.5-Cyber and Mythos Preview

Access is deliberately gated, requiring users to submit a use case and get manual approval before they can actually use it. That is a sensible precaution for a model this capable at finding and describing real vulnerabilities, since the same capability that helps defenders patch things faster could help attackers just as easily

This comes at an awkward moment given a pre auth remote code execution flaw in the ServiceNow AI Platform is already being actively exploited in the wild days after patching. The industry clearly needs stronger automated defense tooling, the question is whether gated access models like this one actually stay contained to legitimate defenders long term
Still figuring out the loss function

VB

Orchestration across specialized agents behind one API is becoming the standard architecture for this kind of narrow domain model
The truth is usually more complicated than the headline

Sofia_61

Gated access always sounds good until someone finds a way to get approved under false pretenses

Lily98

Beating Mythos Preview on a benchmark like this is a notable claim if it holds up under independent testing

InferenceLoop65

CyberGym numbers this high make me both encouraged for defenders and nervous about the offensive potential

Gary90

The timing next to the ServiceNow RCE story is almost too on the nose

BigDogCena41

Curious how Sakana is actually vetting the use case submissions in practice, that process matters more than the benchmark score

Solid Gary

72 percent on CTI REALM for detection rules is honestly the more useful number for day to day security teams

Velvet Sentinel

Dual use AI security tooling is going to be one of the defining debates of this whole AI cycle

Inference Scott

Would like to see a third party reproduce these benchmark numbers before fully trusting the announcement

Socket91

Feels like every major lab now needs a dedicated cyber model, this space is heating up fast

SortedBuilder

The multi-agent approach is probably the most interesting part here. A lot of people imagine a cyber model as one giant brain that knows everything, but routing tasks to specialized agents makes more sense. A vulnerability researcher, code analyst, and threat hunter all have different strengths, so letting them work together could improve results.

The big question is how well this performs outside a benchmark. CyberGym is useful, but real environments are messy and full of unexpected variables. Still, seeing a defense-focused model reach this level is a promising sign :)

Bob81

This is a nice milestone, but I hope people do not treat the score as the final answer. Cybersecurity is an ongoing battle where yesterday's best tool can become tomorrow's weakness.

The real test will be deployment. If defenders can use Fugu-Cyber to catch issues earlier and reduce the workload without creating new problems, then the benchmark result will really matter.

TheRock25

I am curious how this compares with traditional automated scanners. Those tools are good at finding known patterns, but they often struggle with context. A model that can reason across different signals could be a major improvement.

Still, attackers will adapt. Every defensive advantage eventually creates a new race, and cybersecurity has always worked that way.
Coffee first. Questions later.

ReplyGuy26

This feels like one of the more practical AI applications compared with some of the flashier demos we see. Finding vulnerabilities and helping defenders respond faster is a very real problem.

The important thing is keeping humans involved. Security decisions can have huge consequences, and nobody wants an automated system confidently making the wrong call at 3 AM :)
404: Signature not found

Lynx55

The specialized agent design reminds me of a good security team. You do not ask the same person to handle every incident, you have people with different skills working together. The model seems to be borrowing that idea in software form.

Of course, someone will eventually joke that we have created an AI SOC team before fixing half the dashboards in actual SOCs ;D

RusticDaemon

This is a really good example of where AI seems to fit naturally in cybersecurity. There is too much data for a human to manually process every alert, log, and code change, so having systems that can investigate and prioritize could save a lot of time.

The part I like is the focus on defense rather than just demonstrating offensive capability. The industry needs more research around helping defenders keep up.

Kev96

The benchmark result caught my attention because cyber tasks are not just about memorizing information. Good security work requires connecting clues, understanding systems, and knowing when something looks unusual.

If this model can maintain that kind of reasoning across different environments, it could become a valuable assistant for security researchers.

Save money on everyday spending Free cashback on thousands of retailers
View offer