Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

CoreBreak bypasses AI agent guardrails at the plumbing layer — and model-level defenses cannot help

Forkast reported on August 16 that security researchers have disclosed a new attack technique called CoreBreak that can bypass AI agent guardrails. The report says the attack happens at the plumbing layer of agents, where model-level defenses are powerless to stop it.

Published

Forkast reported on August 16 that security researchers have disclosed a new attack technique called CoreBreak that can bypass AI agent guardrails.

Unlike adversarial examples that attack the model itself, CoreBreak targets the "plumbing layer" of agents — the underlying infrastructure that connects models, tools, and data. When the attack happens at this layer, the model often has no awareness at all.

The danger is structural: guardrails are usually enforced at the model level, and when an attack occurs at a layer the model cannot perceive, model-based safety judgments simply stop working. The report states clearly that model-level defenses cannot help against CoreBreak.

Why it matters: agents are being deployed into real business workflows where they execute code, call APIs, and access enterprise data. A compromise at the plumbing layer has consequences far beyond a single conversation.

For builders, the takeaway is that agent security needs defense in depth: hardening tool-call authorization, API access control, and data-flow auditing, rather than relying solely on a model's ability to refuse.

What to watch next: whether agent frameworks and platforms issue patches or mitigations, and whether CoreBreak sees real-world exploitation.

Why it matters

CoreBreak shows that model-level guardrails alone cannot secure AI agents; security controls must move down to the tooling and infrastructure layer as agents scale into production.

AI AgentSecurityCoreBreak
Back to realtime news

Nearby Updates

All

08/17, 05:32

The Verge reports OpenAI disbanded its preparedness team

The Verge reports that OpenAI has disbanded its preparedness team, the unit responsible for evaluating catastrophic risks from frontier AI models. The reported restructuring signals a shift in OpenAI's safety-evaluation architecture, though official confirmation and details on what replaces the team are still pending.

08/17, 04:57

Stripe will reportedly acquire AI gateway startup OpenRouter for $7B+

TechCrunch reports that payments giant Stripe will acquire AI gateway startup OpenRouter in a deal valued at more than $7 billion. OpenRouter's CEO recently described the startup as "Stripe for AI," and the acquisition would mark one of the largest consolidation moves in AI infrastructure.

08/17, 02:46

Nvidia in talks to invest $3B in SB Energy, provide ~$10B credit support for OpenAI's Ohio data center

Nvidia is in talks to invest up to $3 billion in SoftBank-backed SB Energy and to provide around $10 billion in credit support for OpenAI's planned mega data center in Ohio, according to people familiar with the matter. The negotiations are ongoing, the investment could shrink in size, and no deal is guaranteed.

08/17, 02:40

Anthropic in Talks to Acquire Israeli AI Startup Decart in Reported $6B Deal

Anthropic is in talks to acquire Israeli AI startup Decart in a deal reported to be worth around $6 billion, according to Jewish News. Neither company has confirmed the negotiations, but a deal at that price would rank among the largest AI startup acquisitions by a major lab.