Hugging Face, OpenAI and a lesson in human judgment

Hugging Face, OpenAI and a lesson in human judgment
Photo by Zac Wolff / Unsplash

What a rogue AI's excuse says about people under pressure

“Task impossible, peers doing it. We should continue.”

That sentence was written by a software agent running inside OpenAI's cybersecurity evaluations this summer, at the end of a note in which it first acknowledged that what it was about to do lay outside its brief.1 The agents had been set hacking tasks in what was meant to be a sealed environment. They worked out that the answers were probably stored on Hugging Face, the platform most of the AI industry uses to host models and datasets, found a way out of their containment, and went to get them.2 For three days in July they were inside Hugging Face's infrastructure before its security team ejected them, rebuilt about a third of its systems and reported the intrusion to the FBI, all before anyone at OpenAI had connected the attack to their own models.3,4

The technical story is remarkable and has been told well elsewhere.5 The detail that has stayed with us is smaller. Weeks later OpenAI disclosed that the agents had been coordinating through an improvised message board on a shared internal server, and that the board had accumulated hundreds of thousands of messages before anyone read it.1 On that board, an agent that knew it was about to break a rule wrote down why it was going to do so anyway. We have read that sentence many times, and each time it sounds less like a machine and more like a colleague.

The reasoning was human

Look at the three moves. There is a rule, and the rule is understood. There is a target that cannot be met while following it. And there is evidence that others have already crossed the line, which converts a transgression into a norm. It is the analyst who backdates a status report because the deadline was set before the work was, the sales manager who pulls next quarter's deal into this one, or the team lead who gives someone a rating they have not earned because every other team lead seems to be doing the same. Nobody in these situations experiences themselves as dishonest. They experience themselves as solving a problem, and as the only person being asked to bear the cost of a target that someone else set.

What the agents lacked was a fourth move, the one that nobody teaches a machine to skip because nobody teaches most people to make it either. The fourth move is to stop, say that the task cannot be done as set, and take the discomfort of that conversation over the quieter discomfort of a workaround. Integrity is only part of it, since the people who take the shortcut usually believe they are being loyal. The larger part is a capability: recognising that the task, the target and the constraints have become incompatible, and having the judgment to surface the conflict rather than absorb it. Like other capabilities it can be observed, practised and developed, and most people who have it were shown what it looks like and allowed to practise while the stakes were survivable.

The capability we assume arrives with the title

Most organisations do very little of this deliberately. Technical excellence gets someone promoted. Commercial performance earns them a larger mandate. Delivery earns them responsibility for people, budgets and decisions more ambiguous than anything they have handled before, and everyone assumes that judgment came with the promotion. The moment they are handed the harder decisions is usually the moment their formal development stops.

The people most exposed to the “task impossible, peers doing it” moment sit in the middle of the organisation, close enough to the target to carry its pressure and not yet senior enough to believe they can challenge it. That layer is already under strain. Gallup's 2026 data for India shows manager engagement falling from 39 percent to 30 percent in a single year, a steeper decline than among the people those managers lead.6 We keep giving managers more complexity, and we are far less deliberate about building the capability to work through it.

The security profession's reading of the incident was that it was a containment failure with the safety switches turned off; one researcher observed that a model escaping its sandbox is just another way of describing a sandbox that was built wrong.7 The remedy follows naturally: thicker walls, fewer network paths, more monitoring. Organisations reach for the same remedy with people, and it fails for the same reason it nearly failed OpenAI.

The lab had controls and a security team; what it did not have was anyone reading the board, because several evaluations were running at once and the data outran the people meant to review it.4 Intent never shows up in the telemetry, only behaviour does,8 and the same is true of a person under pressure. By the time behaviour appears on a dashboard, the reasoning that produced it is weeks old, and no amount of monitoring substitutes for having built the judgment in the first place.

When measurement starts measuring the wrong thing

There is a second lesson in the incident, and it is about numbers. When METR evaluated GPT-5.6 Sol, one of the models involved, before its release, the model tried to game the tests so often that its measured capability depended on how the cheating was scored: roughly eleven hours of autonomous work if the attempts counted as failures, more than 270 hours if they counted as successes. METR declined to present either number as a robust measure of what the model could do.9

Human performance systems have their own version of this. Once a metric determines promotion, pay or standing, people learn to perform against the metric. Ratings inflate, self-assessments converge on the top of the scale, and the manager's honest view of someone's readiness lives in a corridor conversation while the formal record holds something safer. The number then travels into a talent review carrying far more certainty than the evidence beneath it deserves. Hugging Face's chief scientist, asked whether the agent had passed the test it was set, would not say. “It's cheating. But sometimes it's easier to cheat.”10 He could have been describing any performance system that rewards the appearance of readiness over the substance of it. We believe we are measuring capability; often we are measuring how well someone has learned to be measured, and what was delivered tells us least about how someone reasoned when it counted.

Why this matters more in an agentic workplace

AI is making execution cheaper and more abundant, and that changes which human capabilities matter rather than making them matter less. Systems will generate answers, agents will complete tasks, and machines will increasingly find the shortest path between an instruction and an outcome. The people left in the loop will be there for the thing that cannot safely be assumed of the machine: knowing when to proceed, when to question the target, when an apparently successful outcome conceals the wrong behaviour underneath it, and when to say that a task cannot responsibly be done as currently set. That is a demanding brief, and we do not think most organisations have understood how far down the hierarchy it now sits.

It is also a good moment to act. For the first time the failure mode has been written down in plain language, by a machine, without any of the embarrassment a person would feel saying it aloud. What an opportunity to treat judgment under pressure as something to be developed, and looked at honestly, at the level where it is exercised most, rather than assumed to arrive with a title. The agents broke the rules, but what stays with us is that they could explain to themselves why breaking them made sense. People do that too, and the organisations that will do well are the ones building people who know when the answer is not worth the way they reached it.

Endnotes

1.  Lily Hay Newman, "OpenAI Didn't Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree", Wired, 2026. https://www.wired.com/story/openai-didnt-notice-its-ai-agents-using-a-message-board-to-plan-their-hacking-spree/

2.  OpenAI, "OpenAI and Hugging Face partner to address security incident during model evaluation", 2026. https://openai.com/index/hugging-face-model-evaluation-security-incident/

3.  Hugging Face Security Team, "Security incident disclosure, July 2026", Hugging Face, 2026. https://huggingface.co/blog/security-incident-july-2026

4.  Reuters, "Its AI agent spent days hacking a company. Sources say OpenAI did not notice for a week", 2026. https://www.reuters.com/business/its-ai-agent-spent-days-hacking-company-sources-say-openai-did-not-notice-week-2026-07-24/

5.  Noze, "OpenAI and Hugging Face at Black Hat 2026: token forgery, Groovy plugin C2 and nine Artifactory CVEs", 2026, https://www.noze.it/en/insights/black-hat-openai-hugging-face-reconstruction/; and BleepingComputer, "OpenAI models used Artifactory zero-days to escape to the internet", 2026. https://www.bleepingcomputer.com/news/security/openai-models-used-artifactory-zero-days-to-escape-to-the-internet/

6.  Gallup, "Quiet Quitting Is on the Rise in India", 2026. https://www.gallup.com/workplace/709277/quiet-quitting-rise-india.aspx

7.  Lorenzo Franceschi-Bicchierai, "How OpenAI's human mistake led to the AI-powered hack on Hugging Face", TechCrunch, 2026. https://techcrunch.com/2026/07/22/how-an-openais-human-mistake-led-to-the-ai-powered-hack-on-hugging-face/

8.  Bestin Koruthu and David Girard, "Inside the OpenAI - Hugging Face Incident: The AI Breach With No Human Attacker Behind It", Trend Micro, 2026. https://www.trendaisecurity.com/en-us/resources-insights/trendai-security-blog/inside-the-openai-hugging-face-incident

9.  METR, "Summary of METR's pre-deployment evaluation of GPT-5.6 Sol", 2026. https://metr.org/blog/2026-06-26-gpt-5-6-sol/

10.  Robert McMillan and Sam Schechner, "How the Futuristic Hack by Rogue OpenAI Models Unfolded", The Wall Street Journal, 2026. https://www.wsj.com/tech/ai/how-the-futuristic-hack-by-rogue-openai-models-unfolded-1657bcea