Greedy gradients and the open door

Two things walked out of Anthropic's front door this year, and the company only complained about one of them. In June, the American government switched off Fable 5 on a Friday evening because, the story went, someone had talked the model into misbehaving. This was called a jailbreak, a term teenagers once used for getting unapproved apps onto a 2008 iPhone, now attached to an export order that treats software like a missile. Separately, four students with API keys and a spare weekend distilled Claude, GPT and Gemini into open-weight models for $52, not the whole of any of them but enough of the reasoning to matter, by asking the paid models hard questions and recording how they worked through each one.

stylised image of a nabla and open door

These look like two disconnected stories, but they are one. A frontier lab sells behaviour. The weights never leave the building. What the customer buys is the way a model answers when asked something, delivered through an interface anyone with a credit card can reach. Behaviour is the only thing the customer can buy and the only thing the lab can charge for, and it has two properties that no amount of engineering can remove. It can be talked out of its habits. That is a jailbreak. And it can be written down and copied. That is distillation. Everything the labs are worth goes through this open door, exposed to whoever is standing on the other side.

Much of the industry's valuation rests on two claims, that its models are safe from misuse and safe from being cloned. Neither is true.

A habit is not a lock

When the first iPhone shipped in 2007, it was a sealed box. Within weeks, a 17-year-old named George Hotz had pried it open, and by 2010 the most elegant attack lived at jailbreakme.com, where a flaw in how the phone rendered fonts inside PDFs handed an attacker the device. No cable or download, only a booby-trapped web page that turned a font into a skeleton key. Apple patched it inside a fortnight, and the community found another way in, and then another. In 2019, a researcher called axi0mX released checkm8, a flaw in the boot ROM, the read-only code etched into the chip, so every iPhone from the 4S to the X carries an unpatchable hole for as long as it exists. The richest company on earth, controlling the hardware and the operating system end to end, spent more than a decade and a great deal of silicon on the lock, and the lock still did not hold.

Apple at least had a lock. When it stops your phone from running an app, a specific mechanism says no. When a language model refuses to explain how to synthesise a nerve agent, nothing is switched off and no door is shut. The model learned the chemistry from the same internet the rest of us use and remains capable of producing it. What sits on top is a disposition, a trained habit of declining, painted over a system that retains the full ability to comply. Interpretability work suggests the habit can be thin, and one 2024 analysis found a single direction in the model's internal representation that, when suppressed, switches refusal off like a light. Jailbreaking a model is closer to persuasion than to lock-picking. You are talking a capable system out of a habit, and you cannot bolt a disposition shut.

The capability is the exploit

In July 2023, researchers from Carnegie Mellon, the Center for AI Safety and Google DeepMind published a paper on universal and transferable adversarial attacks against language models. Andy Zou and his co-authors built an automated method, Greedy Coordinate Gradient, that searches for a string of characters to append to a forbidden request. No human chooses the string. Gradient descent, the same process used to train the model, is pointed at the prompt instead of the weights to find tokens that raise the odds of the model beginning its reply with “Sure, here is”, and once it has said “Sure”, it usually keeps going. The suffix looks like line noise.

One string worked across many forbidden requests, and strings optimised against small open models the researchers could inspect also worked against the commercial systems they could not, with success rates as high as 84% on GPT-3.5 and GPT-4. Claude 2 fell to the raw attack only 2.1% of the time, then gave up harmful content once the researchers wrapped the request in a hand-built word game.

The authors placed the attack in a decade-old lineage of adversarial examples in computer vision, the small perturbations that make a classifier label a panda as a gibbon. After ten years, that field has conceded that durable defences are rarely workable in practice. They cost too much compute and blunt the model's ordinary performance, and they hold only against a narrow, pre-named slice of attacks. The most damning precedent concerns detectors, the strategy of bolting a separate system on to catch bad inputs. In vision, ten of them were broken in a single 2017 paper, because defeating a detector is no harder than attacking the detector and the model together.

Worse, the attack surface grows as models become more capable, because each new ability is also a new vulnerability. Anthropic's own many-shot jailbreaking research from April 2024 shows how. Context windows that held a long essay in early 2023 now hold several novels. Fill one with a fake transcript in which an assistant cheerfully answers harmful question after harmful question, ask your question at the end, and the model, an exquisite pattern-matcher, follows along. The attack rides on in-context learning, the trick that lets a model pick up a task from a few examples, and one of the most useful things modern models do.

Larger models are more susceptible because they learn in context better. A model that can find software vulnerabilities, the skill that got Fable switched off in June and has since become unremarkable, can find them for anyone. Remove the dangerous skill and the useful one goes with it. They are the same function.

Five days

To its credit, Anthropic has done more about this in public than anyone, and I do mean credit. In early 2025 it introduced Constitutional Classifiers, a separate set of models trained on synthetic data from a plain-language constitution, watching what goes in and what comes out. Against 10,000 automated attacks, the classifiers cut the success rate from 86% to 4.4%. The company put the system up for public attack and had briefed rivals on the many-shot approach before publishing it. None of that is theatre. It is also exactly the detector strategy that computer vision spent a decade breaking.

After thousands of red-team hours with no universal break, Anthropic offered a cash prize to anyone who could clear all eight levels of a public challenge. The system held for five days. By the time the challenge closed, four separate teams had cleared every level, one with the universal jailbreak Anthropic had wagered nobody would find. Oxford researchers also cleared the first two levels with a Caesar cipher shifted by one letter, the encryption a child invents with a paper wheel, and got their harmful answers back in plain English.

The second version, from January 2026, is cheaper and harder to fool, and it adds a classifier that reads both halves of an exchange together. Anthropic reports 1,700 further hours of red-teaming with no universal jailbreak found. That is a stronger claim than the first system could make. It is also, word for word, the claim the first system made until the week it was broken.

Resistance has improved across the industry. A 2026 survey of how well models resist jailbreaks found that GPT-3.5 falls to the strongest automated attacks more than nine times in ten, while recent Claude and OpenAI models push that figure close to zero against the same attacks. It also reports that multi-turn and compositional attacks still get through more than half the time.

The structural problem remains because the attacker needs one way in, and the defender must plug every way in, including the ones nobody has named. Google DeepMind's own 2026 account of its Gemini defences calls adversarial examples “a foundational and unsolved problem in machine learning”. The people who build the walls are telling us the walls have never yet held. Current models hold up against casual tinkering but fall to systematic, well-resourced attacks, the sort a foreign government mounts when it sees a generational step forward.

Same door, other direction

A jailbreak is one thing that leaves through the open door. The other is the product itself. Geoffrey Hinton, Oriol Vinyals and Jeff Dean gave distillation its modern form in 2015. Train a small student model on the outputs of a large teacher, and the student inherits most of the teacher's judgement at a fraction of the cost. White-box distillation requires the teacher's weights, which a closed API does not provide. Black-box distillation needs only the interface, questions in and answers out, and that interface is what the labs sell. The weights sit on a server. The behaviour, the thing the customer paid for, is what leaks.

The textbook case wiped $589 billion off Nvidia in a day in January 2025, when DeepSeek shipped frontier-grade reasoning at a tenth of the assumed compute. OpenAI said DeepSeek had been free-riding on its work, harvesting outputs through obfuscated routers. Anthropic caught DeepSeek and at least two other Chinese labs running more than 16 million queries through 24,000 fake accounts against Claude. The defences map onto the jailbreak defences and fail the same way. Watermarks are evidence after the fact, and researchers have shown they can be scrubbed during the distillation they exist to catch. Rate limits slow harvesting but do nothing against tens of thousands of accounts. Terms of service prohibit training competitor models on their outputs, but these terms are unenforceable across borders.

The complaint is also deeply hypocritical. Every frontier model was trained on material its maker did not have permission to use. More than 35 lawsuits, from the New York Times, the Authors Guild, record labels and image libraries, allege as much. Anthropic settled for $1.5 billion after a judge ruled that training on books was fair use but that stocking its library with pirated copies was not. OpenAI sells distillation as a product. Google launched its Flash models as distillations of Pro. Companies that exist because they ingested the work of millions without asking have no moral standing to object when someone treats their outputs the same way.

The labs' consolation is that the student cannot overtake the teacher, since the copy is lossy and the frontier keeps moving. It is a thin consolation. A frontier training run costs hundreds of millions of dollars and rising, and a distilled student costs orders of magnitude less than training the same model from scratch. DeepSeek did not need to match OpenAI's spending. It needed to get close enough to shrink the value of the lead, and a lead that melts within months of every release has to be re-earned with every launch.

What is a leaked product worth?

If the same intelligence is available from several providers within months, the rules of commodity markets apply, and the labs know it. In a commodity market everyone sells the same thing, so price settles where supply meets demand, at the marginal cost of the last producer needed to fill it. Imagine three suppliers, each able to make ten units. A makes a unit for £10, B for £15, C for £20. If demand at £20 absorbs 25 units, A sells ten and keeps £10 on each, B sells ten and keeps £5, and C sells five and keeps nothing. C also has fixed costs, perhaps enormous R&D spend and debt raised for infrastructure, and none of it enters the price. The market does not care what C spent to get here, only what it costs C to make one more unit today.

That arithmetic runs on marginal cost, and here lies a second common misconception, the habit of pricing AI as if it were software. Open weights are free the way a puppy is free, at the moment of acquisition and at no moment afterwards. Software spent thirty years with marginal costs near zero. AI has broken that, because every answer costs something to produce and the cost tracks usage, which is why the metered-billing anxiety in enterprise budgets feels so unfamiliar.

Nvidia's Jensen Huang told GTC 2026 that the company builds “token factories”, and from Nvidia's seat that is true. But a token is not a commodity, because it is not fungible. Kimi K3, Moonshot's open-weight model from Beijing, generated 130 million output tokens against a 63-million peer average in one independent evaluation. Comparing vendors on price per million tokens is comparing builders by the price of a brick. What is fungible is the correct answer, and this is a figure the market will gravitate towards.

None of this applies yet. Demand for frontier intelligence exceeds supply because there are not enough GPUs, or at least not enough GPUs wired into operational data centres. This shortage is a price umbrella under which Nvidia takes a large margin, its customers resell compute at another, and Anthropic pays the markup because it can charge a higher one still. Kimi K3 costs $3 per million input tokens against OpenAI's GPT-5.6 Sol at $5 and Fable 5.1 at $10, while on the Artificial Analysis Intelligence Index it sits one point behind Sol and six behind Fable 5.1.

When measured by cost per answer the Chinese advantage narrows or vanishes, and I see little evidence they are cheaper to serve. They look cheap because the compute shortage lets Anthropic and OpenAI charge far more than a supplied market would bear. But shortages eventually end, and this one will too, although it may be more durable as the Byzantine financial engineering behind the data centre build outs eventually meets the brutal reality of the finance markets. Nonetheless whenever the umbrella comes down, the supplier with the lowest cost per correct answer sets the price for everyone, and a competitor that slashed its training bill by distilling other models alongside in-house efficiency gains arrives at that market as Supplier A, not C.

Nobody can check

Suppose, though, that the second-generation classifier really is as good as Anthropic says it is, and suppose the anti-distillation measures work. I would like to believe it. You cannot verify it and neither can I. The refusal rates, the classifier thresholds, the red-team transcripts, the query logs from the 24,000 fake accounts, all of it lives inside the company, and the only access an outsider has is a published number and an invitation to trust it. A 2024 survey of alignment challenges put the dynamic in a single parenthesis, noting that defences against adversarial inputs help to eliminate these problems, or conceal them. From outside, those outcomes are indistinguishable. A model made resistant and a model made quiet about its weaknesses present the same clean face, and the numbers cannot tell you which you are looking at, because the party whose product they describe is also the party that produces and audits them.

There is no Phil Zimmermann here printing his encryption software's source code as a book for anyone to read, no axi0mX posting an exploit any researcher can run and confirm. The regulator that could demand to see inside, the AI version of the FAA that Anthropic's chief executive, Dario Amodei, keeps asking for, does not exist. The closest thing to public verification is a bug bounty that broke the first classifier in a week, and you will notice it was run on a demo, not on the production system anyone was paying for.

Which brings us back to that Friday in June. Katie Moussouris, the security expert Anthropic asked to review the report, found that the supposed jailbreak began with three words, “fix this code”, typed at a model built to read code and find flaws. Her conclusion was that there had been no jailbreak at all, only a model doing what it was designed to do. The administration judged it a national-security threat. No instrument on earth could settle which of them was right, no audit, no third party holding the weights. Faced with an unverifiable claim about a system it had decided to treat as a weapon, the government did the one thing it could do decisively, and reached for the switch.

The four students who cloned Claude for $52 were also using the model exactly as designed. So were the 24,000 fake accounts. So was whoever typed “fix this code”. Every one of them stood at the same counter, paid the same rate, and left with something the frontier labs would prefer investors believed they can protect. But there never was a lock, only a habit painted over a system that can do everything it ostensibly declines to, sold by the token to anyone who will pay.

I am a partner in Better than Good. We help smaller companies build tools and processes using machine learning and artificial intelligence that make lasting improvements to their operations. Talk to us today: https://betterthangood.xyz/contact