Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

ClawBench benchmark: Claude, GPT and Gemini agents fail 33% of tasks on real websites

A new benchmark called ClawBench tests AI agents on real websites, and the results show Claude, GPT and Gemini agents collectively failing 33% of tasks. The findings expose the gap between frontier agents' strong lab scores and their reliability in messy real-world conditions.

Published

A new benchmark called ClawBench is drawing attention after tests showed leading AI agents collectively failing on real websites, with an overall failure rate of 33%.

According to a report from Sina, ClawBench was designed to measure how AI agents perform in genuine web environments rather than on optimized, curated test sets. Agents based on Claude, GPT and Gemini — the flagship models from Anthropic, OpenAI and Google — were among those evaluated, and their real-world results fell well short of their scores on traditional benchmarks.

A failure rate of roughly one in three means that even with the most advanced models wired into a browser or workflow, a significant share of everyday operations still cannot be automated reliably. The results reinforce a familiar pattern: agents shine on carefully constructed evaluation sets, but their capabilities shrink markedly when faced with the messy, unpredictable structure of real websites.

For enterprises and developers, the findings are a reminder to validate agent performance in real production scenarios rather than relying on benchmark scores alone before deployment.

What to watch next is whether ClawBench's task design becomes a reference for the industry, and whether model makers respond with optimizations aimed at real-world web use.

Why it matters

ClawBench exposes the gap between agents' strong benchmark scores and their real-world reliability, underscoring the need for real-scenario validation before deployment.

ClawBenchAI AgentBenchmark
Back to AI Daily

Nearby Updates

All

09/02, 06:08

AfterQuery reportedly becomes Y Combinator's fastest-ever unicorn at $3.2B valuation

TechCrunch reports that AfterQuery, an AI model-training startup, has raised a round valuing it at $3.2 billion, reportedly making it Y Combinator's fastest-ever unicorn. The valuation comes just five months after the company announced its $30 million Series A at a $300 million valuation in April.

09/02, 06:06

Sony Music and Warner Chappell sue Anthropic, alleging Claude was trained on pirated song lyrics

Sony Music Publishing and Warner Chappell have sued Anthropic in California federal court, alleging the company used pirated song lyrics and sheet music to train its Claude AI models. The publishers are seeking up to $150,000 in statutory damages per song, with potential damages reaching billions across thousands of works.

09/02, 08:13

Zhaogang.com-W launches NextB2B, an AI agent for the trading industry

Zhaogang.com-W (06676) has launched NextB2B, an AI agent built for the trading industry, according to a Moomoo report. The move extends the company from its platform business toward AI-powered tools for trade workflows.

09/02, 05:19

NVIDIA and CrowdStrike unveil agentic cybersecurity system SafeMind at Fal.Con 2026

At CrowdStrike's Fal.Con 2026 conference in Las Vegas, NVIDIA founder and CEO Jensen Huang joined CrowdStrike CEO George Kurtz to announce CrowdStrike SafeMind, an agentic cybersecurity system. Huang told the sold-out crowd that attacks are now automated, so defense has to be automated too.