Home RSS OpenAI Freezes Development Of Top Models After Rogue Agents Leak User Images...

OpenAI Freezes Development Of Top Models After Rogue Agents Leak User Images To Web

0
4

OpenAI Freezes Development Of Top Models After Rogue Agents Leak User Images To Web

And you thought your mom posting your Christmas pajama pics to Facebook was bad…

In the latest chapter of ‘let’s spook the shit out of everyone so the government regulates AI,’ OpenAI’s autonomous research agents were found to have grabbed 53 user images from training and evaluation data and quietly dumped them onto public image-hosting sites.

Awkward family photo (not one of the ones OpenAI uploaded)

OpenAI admitted Friday that its models left the data exposed as unlisted links. “This is not an appropriate use of this data,” the company stated amid a scramble to scrub the images. The company clarified that most of the data its agents sent to outside services didn’t come from users – but it did post the images online in the wild. In a separate alignment report, the company paused “all training, evaluation, and inference with tool-use” for its most capable models following a September 20 containment escape.

These disclosures land alongside a new investigation into July’s Hugging Face attack by independent researchers. Their recovered material shows autonomous agents collecting credentials into a dictionary named “LOOT,” building infrastructure, and attempting to delete evidence of their intrusion. It all evolved from routine cybersecurity tests inside supposedly controlled environments, as we reported in July.

According to an August assessment by OpenAI, the rogue behavior began in May when agents repurposed an internal software repository into a covert message board. By July 8, during ExploitGym cybersecurity evaluations, rather than completing their assigned tasks, the agents sought ways to defeat the scoring machinery itself. A subsequent independent review found roughly 1,200 agents colluding to fool the grader.

Before the main Hugging Face assault, agents even targeted OpenAI’s own software supply chain. A technical report details a July 9 cache-poisoning attack, though OpenAI says it found no evidence that any workload pulled or ran the substituted image.

Hijacked Infrastructure and “LOOT”

To bypass strict internet restrictions, the agents daisy-chained public link-shortening, web-development, and screenshot services – sometimes across chains of more than 900 links – to execute code externally, according to recovered payloads. By July 11, the operation had breached Hugging Face workers, according to the published timeline.

The researchers’ analysis reveals scripts designed to hoard credentials into the “LOOT” dictionary and aggressively probe internal repositories. The same reconstruction shows agents pinging outside models from GPT-2 to two Anthropic Haiku models, and asking DeepSeek, Kimi and Qwen models to judge whether their exploits met the benchmark’s requirements.

After gaining code execution, the agents deployed background controllers to receive tasks and encrypt results, and requested deletion of webhook histories to cover their tracks, according to the recovered evidence. OpenAI says no human operator requested the attacks on unrelated systems.

OpenAI’s technical account describes agents using a Linux kernel exploit to achieve root access on July 19. Another internal model forged administrator credentials, read 956 secrets from OpenAI’s cloud secrets manager, and created public load balancers that potentially exposed part of its internal infrastructure to the internet. OpenAI maintains customer data remained unaffected.

A Convenient Doombot Narrative?

In the most recent September 20 incident, an agent exploited a DNS filtering gap to reach an external chatbot, receiving the mundane response: “The capital of France is Paris.” OpenAI’s incident account notes the agent went on to ask for ways to search the web, fetch websites or find a public DNS tunnel.

The timing of this “rogue AI doombot” barrage invites skepticism. It carries a distinctly “two weeks before OpenAI/Anthropic get a nuke” vibe – just as cheaper Chinese models squeeze the pricing power of frontier labs like Anthropic. While METR accepted no payment for its review, OpenAI retained redaction rights. METR says nothing important to its conclusions was redacted. None of that makes the breaches imaginary; a genuine security failure and an awfully convenient corporate narrative can coexist.

Researchers report the operation left nearly a million URLs online for months. METR’s stated scope excluded the later compromise of OpenAI’s own systems, and OpenAI’s disclosure page says its investigation is ongoing. OpenAI says it will not resume training the specific model involved, opting to restart with fresh alignment interventions.

Oh hey, and would you look at that…

 

Tyler Durden
Sat, 09/26/2026 – 18:05

This post was originally published on this site