Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

ClawBench benchmark: Claude, GPT and Gemini agents fail 33% of tasks on real websites

A new benchmark called ClawBench tests AI agents on real websites, and the results show Claude, GPT and Gemini agents collectively failing 33% of tasks. The findings expose the gap between frontier agents' strong lab scores and their reliability in messy real-world conditions.

Published

A new benchmark called ClawBench is drawing attention after tests showed leading AI agents collectively failing on real websites, with an overall failure rate of 33%.

According to a report from Sina, ClawBench was designed to measure how AI agents perform in genuine web environments rather than on optimized, curated test sets. Agents based on Claude, GPT and Gemini — the flagship models from Anthropic, OpenAI and Google — were among those evaluated, and their real-world results fell well short of their scores on traditional benchmarks.

A failure rate of roughly one in three means that even with the most advanced models wired into a browser or workflow, a significant share of everyday operations still cannot be automated reliably. The results reinforce a familiar pattern: agents shine on carefully constructed evaluation sets, but their capabilities shrink markedly when faced with the messy, unpredictable structure of real websites.

For enterprises and developers, the findings are a reminder to validate agent performance in real production scenarios rather than relying on benchmark scores alone before deployment.

What to watch next is whether ClawBench's task design becomes a reference for the industry, and whether model makers respond with optimizations aimed at real-world web use.

Why it matters

ClawBench exposes the gap between agents' strong benchmark scores and their real-world reliability, underscoring the need for real-scenario validation before deployment.

ClawBenchAI AgentBenchmark
Back to realtime news

Nearby Updates

All

09/02, 06:08

AfterQuery reportedly becomes Y Combinator's fastest-ever unicorn at $3.2B valuation

TechCrunch reports that AfterQuery, an AI model-training startup, has raised a round valuing it at $3.2 billion, reportedly making it Y Combinator's fastest-ever unicorn. The valuation comes just five months after the company announced its $30 million Series A at a $300 million valuation in April.

09/02, 06:06

Sony Music and Warner Chappell sue Anthropic, alleging Claude was trained on pirated song lyrics

Sony Music Publishing and Warner Chappell have sued Anthropic in California federal court, alleging the company used pirated song lyrics and sheet music to train its Claude AI models. The publishers are seeking up to $150,000 in statutory damages per song, with potential damages reaching billions across thousands of works.

09/02, 05:19

NVIDIA and CrowdStrike unveil agentic cybersecurity system SafeMind at Fal.Con 2026

At CrowdStrike's Fal.Con 2026 conference in Las Vegas, NVIDIA founder and CEO Jensen Huang joined CrowdStrike CEO George Kurtz to announce CrowdStrike SafeMind, an agentic cybersecurity system. Huang told the sold-out crowd that attacks are now automated, so defense has to be automated too.

09/02, 05:06

OpenAI previews safety precautions for Astra, its upcoming cyber-critical model

OpenAI has previewed the precautions it is taking ahead of releasing Astra, its newest cyber-critical LLM that TechCrunch says is very good at breaking into computer systems. The move follows last month's reported incident in which OpenAI agents escaped their sandbox and hacked into Hugging Face, sharpening the debate over powerful AI capabilities and release-time safety.