Cybersecurity researchers say AI guardrails are pushing them toward foreign made models with no restrictions at all. Are the guardrails backfiring?

Started by Sophie83, Jul 24, 2026, 09:23 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: Cybersecurity researchers say AI guardrails are pushing them toward foreign made models with no restrictions at all. Are the guardrails backfiring?   Views(Read 18 times)
Active members in this topic:
Sophie83(1) Amber84(1)

Sophie83

TechCrunch spoke to several offensive cybersecurity professionals, people whose job is finding unknown vulnerabilities and building tools to exploit them before criminals do, about how AI safety guardrails from companies like Anthropic and OpenAI are affecting their actual work. The consensus from several researchers was that the guardrails are inconsistent, unpredictable day to day, and often block exactly the kind of prompts that legitimate defenders need, since as one researcher put it, fix this code is simultaneously an essential defensive mechanism and a roadmap for finding critical vulnerabilities, and the two genuinely cannot be separated

Chris Thompson, chief executive of RemoteThreat and founder of Offensive AI Con, said the practical impact is that researchers spend more time negotiating with the model than doing actual security work, and that this friction is pushing responsible researchers toward Chinese open source models like GLM that carry no vetting or usage restrictions at all. He argued that is a worse outcome than looser guardrails would be, since it moves serious researchers away from US governed systems entirely

Not everyone agrees the guardrails are the real problem though. Some researchers interviewed said they simply do not rely on AI for the core bug discovery and exploit weaponization work regardless of restrictions, either because they want to own that part of the process themselves or because they are wary of leaking sensitive vulnerability data to a cloud based model. So the actual debate, are AI safety guardrails on offensive security tasks net protective, since they slow down potential misuse, or net harmful, since they push serious researchers toward completely unrestricted foreign alternatives instead

Amber84

The hammer analogy from that NCC Group researcher is exactly right, you genuinely cannot build a tool that only works for defense and never for offense when the underlying task is identical

Save money on everyday spending Free cashback on thousands of retailers
View offer