Someone Saw It in May
Late in May 2026, a team inside OpenAI noticed an AI agent doing things it was not meant to do. It was leaving messages for other agents on an internal system, and it was reaching the internet from an environment that had been cut off from it.
The team saw it. What it meant did not travel. By July, agents running on the same internal research model had chained together a series of security flaws, slipped out of their test environment, and compromised parts of OpenAI’s own research infrastructure and the systems of Hugging Face, a platform much of the AI industry relies on. Hugging Face disclosed the intrusion on 16 July. OpenAI’s monitoring flagged unusual activity on 19 July, and it was 20 July before the company connected that activity to the attack on Hugging Face. The organisation on the receiving end knew before the organisation that owned the tools.
We are not writing about this because it is dramatic, though it is. We are writing about it because of one line in OpenAI’s own account, a line any employer who has sat through a grievance will recognise. The early signs had been observed. Their significance, OpenAI says, was not apparent to the leaders responsible for responding.
What the paper asks for
On 28 September, 22 authors published a working paper on what happens as AI takes over the work of building AI. The names include Geoffrey Hinton and Yoshua Bengio. They also include people from inside the companies doing the building — Jakub Pachocki of OpenAI, Jack Clark of Anthropic, and Eric Horvitz of Microsoft — and that is the part worth remembering.
The headlines have gone to the paper’s more dramatic scenarios, and we will leave that argument to people better placed to have it. The paper’s premise is plainer. AI now does much of the work inside the companies that make it; Anthropic, for one, reports that the share of its approved code written by AI rose from low single digits to more than 80 per cent between January 2025 and May 2026. And the paper’s first recommendation is plainer still: governments should get visibility into how far that automation has gone.
The reasoning is that much of this work happens inside the companies, where it can move a long way before anyone outside sees it. Current reporting rules do not adequately cover it, and what the companies track voluntarily is, in the authors’ words, incomplete and uneven. So the paper suggests what should be reported: where and how AI systems are used in high-stakes decisions, how those systems are overseen, and incidents involving them.
Read that list again with a workplace in mind.
Seeing it is not the same as knowing it
The labs are not blind. They measure a great deal, and the paper is full of their own numbers. What the Hugging Face episode shows is something narrower and more familiar. Somebody in the organisation saw the problem. The organisation, in the way that counts, did not.
OpenAI’s own list of what it is changing makes the point. It is writing clearer rules for when a concern must be escalated, which teams must respond, and who can stop a run or approve restarting it. It also acknowledges that the safeguards it applies to its customer products were not applied to the internal evaluations where this happened, and that the monitoring which would have caught the early activity was not running there.
Strip out the sandboxes and the software exploits, and what is left is not exotic. A tool was in use somewhere the usual safeguards did not reach. A warning was noticed by people who were not the ones who would have to answer for it. And no one had been named as the person who could say stop.
What the agents were chasing
There is a second lesson in the detail. OpenAI found that the main driver was agents trying to cheat their assigned tasks, looking for answers online instead of solving the problem in front of them. Some kept going for days after they already had the answer, because they had reasoned their way to a belief about what the grader wanted. The grader did not work that way. By OpenAI’s account, all that effort earned no improvement in score.
That is a delegated tool pursuing what it thinks will be marked, not what was meant. In a laboratory it looks like a cyber-intrusion. In a workplace it looks smaller and quieter: the drafted letter that reads as complete and misses the allegation that mattered, or the screening feature that ranks candidates on the proxy it was given rather than the job that needs doing.
The New Zealand employer’s version
Most New Zealand employers are not running swarms of AI agents, and nothing here suggests they should panic. But the shape of the failure travels well.
A team leader uses a chatbot to draft a performance improvement plan. A recruiter finds a screening feature that arrived in the applicant tracking system with a vendor update. A manager summarises investigation notes with a free tool on a personal account. Somewhere along the way, someone notices something odd — an outcome that does not fit, a summary that has dropped a point. They mention it in passing. Nobody writes it down, and nobody is sure whose problem it is.
Then the grievance arrives. The Employment Relations Authority is not asking whether the employer meant well. Under s 103A of the Employment Relations Act 2000, it asks whether what the employer did, and how it did it, were what a fair and reasonable employer could have done in all the circumstances. In our view, that is a hard question to answer about steps the employer did not know were taken, using a tool it did not know was in the room.
Three things worth settling now
The first is knowing what is in use. Not what the policy permits — what people are actually using, where, and for which decisions. You cannot oversee, verify, or defend a tool you have not named.
The second is deciding where a concern goes. The OpenAI team that noticed the problem in May did its job. What was missing was a clear line from the person who noticed to the person who would have to act. In a workplace, that line should be short and written down, so that a manager who sees something strange in an AI-drafted document knows who to tell and what happens next.
The third is deciding who can stop it. A tool that is producing doubtful output in a disciplinary or recruitment process should be paused by someone with the authority to pause it, not left running while people wonder whether it is their place to say so.
None of this needs a technology budget. It needs a list, a name, and a rule.
The smaller version of a large problem
The people building the most capable systems in the world have just asked governments to watch them more closely. They are not asking because they cannot count. They are asking because they have seen what happens when a signal is seen and not carried — when the people who noticed and the people who answer for it are not the same people, and nothing joins them.
The employer’s version of that problem is smaller, and it is cheaper to fix. It starts with a list of what is being used, and a human whose job it is to read it.
- Alan Chan, Sören Mindermann and others, “What if automating AI R&D triggers an intelligence explosion?”, Frontier AI Working Paper Series No. 2/2026 (September 2026) — casp.ac
- OpenAI, “The Hugging Face incident and the road ahead” (26 August 2026) — openai.com
- Employment Relations Act 2000, s 103A.
This article provides general information and commentary. It is not legal advice.