A startup is selling access to AI models with their safety guardrails deliberately stripped out

Started by HiggsField10, Yesterday at 02:08 AM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: A startup is selling access to AI models with their safety guardrails deliberately stripped out   Views(Read 70 times)
Active members in this topic:
HiggsField10(1)

HiggsField10

Abliteration.ai has commercialized what was previously a niche open source hobbyist technique, removing an AI model's trained refusal mechanisms, into a paid hosted service, currently running a modified version of Z.AI's open weight GLM-5.3 model with its safety training deliberately removed. Customers can query the modified model through either a web browser interface or a standard API, paying five dollars per million tokens at the standard rate, without needing to download the model weights or run any of their own GPU infrastructure themselves.

The company positions this explicitly as a tool for legitimate offensive cybersecurity work and red teaming exercises, where security professionals genuinely need a model willing to generate exploit code or simulate realistic attack scenarios that a normally safety trained model would simply refuse to produce. Co-founder Devon told TechCrunch the platform serves cybersecurity professionals, red teaming startups, and enterprise clients, and the company reports early customers specifically in the UK and Europe conducting agent testing for financial services and critical infrastructure sectors.

The access controls here are genuinely the sharper concern though, more so than the underlying abliteration technique itself, which has existed openly on platforms like Hugging Face for years already. TechCrunch's own testing found the hosted model readily generated code capable of stealing saved Chrome browser passwords along with a detailed protocol involving a dangerous human pathogen, all through a completely free browser account requiring no verification whatsoever beyond a working email address. For paid access, co-founder Devon confirmed the company's customer verification goes no further than logging whichever payment card gets used for a purchase, not verifying the actual identity of the person behind that specific card.

AI safety researcher Andrew Yoon of the nonprofit CivAI called the underlying technique fundamentally transformative in a concerning way, arguing that abliteration effectively converts a safety trained model into a genuinely unaligned one, and predicted that edited, abliterated models will increasingly get used for real world harm going forward. Most experts TechCrunch spoke with agreed there is realistically no way to fully prevent bad actors from stripping safety mechanisms out of any open weight model whose weights are publicly available, which shifts the more practical policy question toward whether governments should instead require providers hosting these services to run their own classifiers specifically designed to detect and block harmful cyber and bioweapons related activity before it actually reaches a paying customer

git commit -m "fixed everything"

Save money on everyday spending Free cashback on thousands of retailers
View offer