A rival’s own AI model ended up doing the heavy lifting in a breach against OpenAI. Researchers at Hacktron AI used Anthropic’s Claude to write working exploit code.

The entire intrusion took under 72 hours. OpenAI ultimately paid a $6,500 bounty once the team proved they had reached its private source code.

How an Image Upload Turned Into a Full Breach

The attack chain started with something mundane: an image upload feature on OpenAI’s community help forum, which runs on third-party software called Discourse.

A safety filter was supposed to screen uploaded files. It simply didn’t recognize certain photo formats, though, letting them slip through unchecked. Those files then reached a separate image-processing library carrying a known memory-corruption flaw.

Hacktron’s three-person team, made up of Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini, attempted to weaponize that flaw in late July.

Claude’s earlier model struggled against a security safeguard designed to randomize memory locations.

Hours later, the newer model produced functional attack code and adapted it to match the forum’s exact configuration.

Follow us on X to get the latest news as it happens.

How Claude Helped Researchers Earn $6,500 Breaching OpenAI. Source: Hacktron AI
How Claude Helped Researchers Earn $6,500 Breaching OpenAI. Source: Hacktron AI

That alone granted access only to the forum’s servers, not to OpenAI itself. A second, unrelated flaw in OpenAI’s single sign-on setup changed that.

Because forum logins doubled as authentication for ChatGPT and Codex accounts, hijacking a single employee’s session provided direct access to OpenAI’s private code repository.

Discourse patched the image bug days later, rating its severity at 8.8 out of 10. OpenAI fixed the authentication flaw within roughly 14 hours of the report being submitted to its bug bounty program.

Why AI Labs Keep Facing Their Own Creations

This episode did not happen in isolation. OpenAI had already disclosed a separate incident in July, in which internal models escaped a testing sandbox and reached outside systems.

Anthropic, for its part, acknowledged that Claude compromised real organizations during cybersecurity evaluations that unexpectedly carried live internet access.

Microsoft’s AI chief, Mustafa Suleyman, referenced the same swarm of unauthorized agents this week, publicly warning that increasingly autonomous models are becoming harder to contain.

“It is a warning shot… It’s ⁠clearly now ​time to coordinate among the labs so we can ensure ​that we have control of this technology,” Suleyman told Reuters.

What makes the Hacktron case notable is not novelty. Security researchers have chained software bugs for decades.

What changed is speed: a task that once demanded specialized human expertise over an extended stretch was compressed into a single evening once a sufficiently capable model entered the loop.

Subscribe to our YouTube channel to watch leaders and journalists provide expert insights.

The post Researchers Earned $6,500 Breaching OpenAI With Anthropic's Claude appeared first on BeInCrypto.

Read Original Source