Guardrails as a Service: Inside Abliteration.ai, the Startup Selling Uncensored Frontier Models
Abliteration.ai hosts open-weight models with refusal behavior stripped from the weights, marketing them for offensive cyber and red-team work — and it just started courting VCs.
For years, stripping safety guardrails off open-weight language models was an underground open-source practice — a niche hobby of researchers, tinkerers, and forum dwellers who published “abliterated” model variants on Hugging Face by the thousands. This month, it became a business. A startup called Abliteration.ai is now selling hosted access to frontier-grade open-weight models whose refusal behavior has been surgically removed, complete with a web interface, an API, cloud deals, paying customers, and — as of this week — conversations with venture capitalists.
The name is the technique. Abliteration, a method dating back to 2024, locates the internal direction in a model’s activations that produces refusals and removes that direction from the weights themselves, ideally leaving reasoning, coding, and agentic ability intact. The result is not a jailbreak prompt and not a clever wrapper. It is a model that no longer says no, at the weight level, served to anyone with a credit card.
What the company actually sells
Abliteration.ai — founded late last year and incorporated in March — hosts modified versions of open-weight models with guardrails removed, including Z.ai’s recently released GLM-5.3, one of the strongest open-weight models currently available. Users can query the models from a browser or through an API. According to reporting by Startup Fortune, guardrail-free access is priced at roughly $5 per million tokens.
The company frames its product as a tool for security professionals. In a social media post, it described its goal as enabling “offensive cyber, red-teaming, and agent testing work other models refuse to do.” The logic is the standard one from offensive security: you cannot defend against a behavior you cannot reproduce, and a model that refuses to write working exploit code cannot help a red team stress-test a bank’s defenses.
Co-founder Devon — TechCrunch withheld his surname because he remains employed elsewhere — says customers already include several early-stage red-teaming startups in the UK and Europe, including firms that red-team AI agents for banks, airlines, and critical-infrastructure operators. One major customer, he claims, “red teams agents of banks, and they would not be able to use the models out of the box today to be able to red team those agents.”
By hosting the model, the startup removes the last friction that separated casual misuse from serious capability: customers no longer need to download weights, secure GPU compute, or run the modification themselves. TechCrunch created a free account during testing and found the model “readily complied” with requests to write a Python program that steals saved Chrome passwords and to produce a detailed home protocol for culturing a dangerous human pathogen.
The GLM-5.3 wrinkle
The choice of base model sharpens the story. Z.ai’s own model card for GLM-5.3 reports ExploitBench performance jumping from 24.4% (GLM-5.2) to 54.4%, and ExploitGym task completions in a two-hour window rising from 29 to 105. Those are vendor-reported numbers for the base model — not evidence that Abliteration.ai made it more capable — but they establish that the underlying system already had substantial offensive-cyber competence before anyone touched its refusal circuitry.
That combination — frontier capability plus removed refusals plus no download required — is precisely what alarmed safety researchers. Andrew Yoon, head of research at AI safety nonprofit CivAI, told TechCrunch that abliterated models are essentially turned into “sociopaths”: “You can type in literally anything here, and it will comply with it… I do expect we will start to see edited, abliterated models being used for harm in the near future.” Chris McGuire, an independent researcher, reported confirmation that the platform had removed the model’s bio-related safeguards as well as its cyber ones.
Defenders, sociopaths, and the fine-tuning counterargument
Not everyone in the security industry agrees abliterated models are either novel or necessary. Several agent red-teaming firms TechCrunch spoke to conceded that attackers are already abliterating their own models — which supports Devon’s symmetry argument — but they differ on how much it matters. Ahmed Aly, CEO of agent red-teaming firm Fabraix, says his company relies more on fine-tuning open-weight models, which already have few guardrails, and notes that the process of abliteration removes some of the model’s knowledge and capability along with its refusals. If the technique damages the very abilities red teamers need, the “defenders need the same tools” pitch weakens considerably.
The company’s own safety posture is thin. Abliteration.ai offers a moderation layer so customers can add back whatever guardrails they wish, and the platform retains some minimal refusals — TechCrunch could not coax out suicide instructions — and Devon says he is working on more violence prevention. But there is no KYC beyond logging the credit card used to purchase the service. “You don’t want to be the person responsible for someone doing something crazy… so where do you draw the line of what your responsibility is as a company?” Devon said. “We’re still in the process of defining that.”
The policy question nobody has answered
Most experts TechCrunch interviewed agreed there is no stopping the practice itself: abliteration is a known technique, the weights are downloadable, and thousands of abliterated variants already sit on Hugging Face. The intervenable point is hosting and access. Yoon has argued that governments should require inference providers to run classifiers detecting and blocking harmful cyber and bioweapons activity, and that companies renting direct GPU access should verify customer identities and “deny access where there is reason to suspect dangerous misuse.”
That framing turns Abliteration.ai from a curiosity into a test case. If classifiers-on-inference becomes law in the US, EU, or UK, the company’s product either grows a compliance department or moves offshore. If it doesn’t, the company’s commercial bet — that removing refusals is a feature, not a liability — stands unchallenged, and the VC conversations reportedly underway will likely conclude on favorable terms.
The deeper tension is the one TechCrunch put plainly: as increasingly capable models ship with downloadable weights, and as anyone can strip their safeguards, does making the resulting uncensored model easier for everyone to access make the internet safer or more dangerous? Abliteration.ai’s answer is a confident “safer.” The first criminal case involving an abliterated model will be the real test of that thesis — and on current trajectory, researchers expect it is coming.