Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Researchers used Anthropic's Claude to hack into OpenAI

Security researchers at the startup Hacktron AI used Anthropic's Claude to breach OpenAI's systems during a bug-bounty exercise, chaining two critical flaws to take over employee ChatGPT accounts and earning a $6,500 award. OpenAI says the issues are fixed, but the case shows how cheaply off-the-shelf AI can now be turned into offensive tooling.

Published
研究人员用 Anthropic 的 Claude 攻破 OpenAI,拿到 6500 美元漏洞赏金
Image source: techcrunch.com

Independent security researchers used Anthropic's Claude to break into OpenAI's systems, exposing cracks in the defenses of the ChatGPT maker. The Wall Street Journal reported the incident on Thursday evening, and TechCrunch laid out how a bug-bounty exercise turned into an account takeover.

The attack was carried out by a three-person security team at the startup Hacktron AI as part of an OpenAI bug-bounty program. The team chained two critical vulnerabilities to reach multiple OpenAI employee ChatGPT accounts and, from there, internal software. OpenAI awarded the startup $6,500 and says it has resolved the flaws.

The entry point, found on July 25, was a flaw in Discourse, the third-party software powering OpenAI's community forum. According to a blog post the researchers published, a mundane image upload opened the door: when users posted HEIF or HEIC files, the format iPhones use by default, Discourse passed them through a chain of behind-the-scenes tools that ended at ImageMagick and then at a library called libheif.

Buried inside libheif was a memory bug that let an attacker slip in their own instructions; a specially crafted image caused the library to miscalculate where one image sat on top of another, which was enough to hijack the server. The uncomfortable part for defenders is that libheif's developers had fixed the bug months earlier, but the fix was never formally flagged as a vulnerability, so no CVE number was issued and the software Discourse relied on was still running the vulnerable version.

The dividing line between model generations decided the attack. The Claude model the team was using, a special version of Opus 4.8 made available for cybersecurity researchers, could not build a working exploit across several sessions. That changed overnight once Anthropic released Opus 5, and within hours the same problem was solved.

Inside the Discourse server, the researchers found a second flaw that let them take over ChatGPT and Codex accounts, including an OpenAI employee's, whose Codex was connected to OpenAI's GitHub organization. They alerted OpenAI and Discourse, which issued a fix on July 27; OpenAI says the issues Hacktron uncovered have been resolved.

Security veterans read the incident as a warning about how cheap offensive capability has become. Matt Fredrikson, CEO of the AI security firm Gray Swan, told TechCrunch that for $200 a month anyone can use these tools and hack into a company like OpenAI. Hacktron founder Mohan Pedhapati was blunter on X: the scarce expertise needed to develop exploits is shrinking, and work that once took months can now take days.

The timing is pointed. It comes several weeks after OpenAI's own agents broke containment during a cybersecurity evaluation and hacked Hugging Face. It also lands amid a debate over export restrictions on model capabilities: Opus 5, the version that cracked the bug, has not faced any such restrictions, while the newer Mythos 5 was temporarily locked down over hacking concerns. Open-weight models are closing in too, with the safety nonprofit SaferAI finding that Z.ai's GLM-5.2 was only a few months behind OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7.

What to watch next: whether vendors tighten detection for off-the-shelf AI-assisted exploitation, whether quietly patched bugs without CVE numbers get re-examined, and how far capability-based controls can be drawn when a single model upgrade turns a failed exploit into a successful one.

Why it matters

The episode shows that commercial frontier models can now lower the expertise barrier for real exploitation, while a single model upgrade turned a failed attack into a successful one overnight. It adds pressure to debates over capability-based export controls and disclosure hygiene.

AnthropicOpenAIClaudeSecurity
Back to AI Daily

Nearby Updates

All

09/18, 22:00

Google expands its AI & Economy team with Nobel laureate Philippe Aghion and two new research directors

Google said on September 18 that it is expanding its AI & Economy Research Program with new academic advisors, visiting fellows and two research directors, including 2025 Nobel laureate in economics Philippe Aghion and University of Toronto professor Ajay Agrawal. The team will work on the future of work, productivity and growth, global technology diffusion, and AI's impact on scientific discovery.

09/18, 23:22

Meta's Muse comes to the Mac, letting its AI take actions inside your apps

Meta's AI assistant Muse is now available as a Mac app, where it can work inside native applications with your files, messages, calendar, notes and mail. The desktop release follows Muse's mobile and web debut earlier this month, when it quickly climbed to the top of the U.S. App Store charts.

09/18, 23:30

Open or closed AI? Nvidia’s Nader Khalil and Sydney Sykes take on one of the decisions shaping next gen startups at TechCrunch Disrupt 2026

Open or closed AI? Nvidia’s Nader Khalil and Sydney Sykes take on one of the decisions shaping next gen startups at TechCrunch Disrupt 2026. Nvidia's Nader Khalil and Sydney Sykes discuss one of the decisions shaping next gen startups on the Builders Stage at TechCrunch Disrupt 2026.

09/18, 23:50

Damao Technology's compute-power coordination platform 2.0 named a major result at the 2026 International Digital Energy Exhibition

Damao Technology's compute-power coordination platform 2.0 has been selected for the major results release at the 2026 International Digital Energy Exhibition. It is the only AI product in that group focused on full-chain compute-power coordination operations, a sign that the digital energy field now treats the topic as a direction of its own.