Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Vals, backed by Andreessen Horowitz, wants to be the gold standard in AI benchmarking

Vals, a startup backed by Andreessen Horowitz, is trying to become the gold standard for AI benchmarking by offering a neutral and trustworthy alternative to the leaderboards flooding the AI world. The company is betting that as model releases multiply, neutral third-party evaluation becomes a decision-making tool rather than marketing material.

Published
a16z投资的Vals:想在AI基准测试里做成“黄金标准”
Image source: techcrunch.com

TechCrunch reported on September 19 that Vals, a startup backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking. The company is positioning itself as a neutral and trustworthy resource in a world increasingly inundated by AI models.

The problem it targets is familiar to anyone following the field, but it keeps getting sharper: the more models there are, the messier evaluation becomes. Vendors publish scores that favor their own systems, and the sources of the numbers and the way they are scored are often opaque, making side-by-side comparison on real tasks difficult.

Vals' answer is to fill that gap. As TechCrunch describes it, the company wants to make AI benchmarking a more neutral and trustworthy resource, so that evaluation stops serving mainly as launch-day marketing and becomes a reference that companies and technical teams can actually use.

The value of that kind of service shows up in decisions. When enterprises pick a model, buy access to one, or wire it into a production workflow, they want reproducible, comparable third-party data rather than a vendor's own demo. Reliable benchmarking is turning into infrastructure for the AI supply chain.

The backing from Andreessen Horowitz adds a venture-capital dimension to the story. Benchmarking is not a model, and it does not sell compute directly, but it carries influence over judgments about which systems are strongest, putting it at a key point between model makers and their customers.

The open question is whether investment can coexist with neutrality. Vals has to keep the trust of model makers, developers and paying customers at the same time, while defending its methods against the familiar charges of bias and score gaming.

What to watch next is adoption: whether a shared evaluation standard from Vals gets widely used across the industry, or whether third-party benchmarking keeps getting diluted by the in-house tests each lab publishes for itself.

Why it matters

Third-party benchmarking is becoming infrastructure for model selection, and whoever makes those results credible gains influence over the industry's consensus on which systems are best.

ValsBenchmarkAndreessen Horowitz
Back to realtime news

Nearby Updates

All

09/19, 19:43

A 27B model that builds web pages in minutes: Qwen 3.8 put to the test

A hands-on test by QbitAI shows Qwen 3.8 27B turning a single prompt into a working data-analysis tool in about five minutes, and into a convincing lookalike of a 12306 train-ticket page. The same test also exposes the limit: the generated pages connect to no live data, accounts, payments or backend, so they impress as interfaces but cannot complete real tasks.

09/19, 16:28

Huawei's Wang Tao: Building the AI compute foundation takes far more than one good chip

At this year's Huawei Connect, rotating chairman Wang Tao unveiled the Ascend 960 and its 960 supernode along with a one-generation-per-year roadmap running to Ascend 970 in 2028 and Ascend 980 in 2029. Huawei also introduced the industry's first supernode built on near-package optics, and said its CANN software ecosystem has crossed an inflection point as it positions itself as a compute-foundation supplier.

09/19, 14:44

Terence Tao Launches SAIR Foundation's Open Math Model Initiative

Fields Medalist Terence Tao has announced the launch of SAIR Foundation's Open Math Model initiative, which aims to bring academia and industry together to build open-weight models and open-source tools for mathematics and scientific research. The first phase targets everyday research work such as understanding arguments, checking literature, writing code, and formalizing proofs, under principles of open licensing, reproducible evaluation, and community governance.

09/19, 08:12

Tilly Norwood's AI Press Tour Stumbles as Synthetic Actor Glitches Mid-Interview

TechCrunch AI reports that Tilly Norwood, an AI performer, is midway through a press tour that is going about as well as one might expect for a synthetic celebrity. In one particularly odd interview, Norwood appeared to malfunction and started speaking in Chinese.