Edited by Editor-in-Chief, The Indus Pulse 19 Sept 2026, 10:29 PM 2 min readai

Vals Secures $40 Million to Standardize AI Model Benchmarking

San Francisco-based startup Vals has raised $40 million in a Series A funding round led by Andreessen Horowitz, aiming to establish a neutral standard for evaluating artificial intelligence models. The company, founded in 2024 by Rayan Krishnan, seeks to address a growing credibility crisis in the AI industry where developers often rely on cherry-picked metrics and legacy benchmarks that fail to capture real-world performance.

Addressing the Benchmarking Crisis

As the number of language models proliferates, the industry has struggled with fragmented evaluation methods. While major players like OpenAI, Google, and Anthropic frequently publish performance data, these metrics are often optimized for specific tests, leading to a disconnect between marketing claims and actual utility in production environments. Vals positions itself as a neutral arbiter, similar to a consumer advocacy organization, by conducting private evaluations that assess how models perform on complex, industry-specific tasks including coding, finance, and law.
Unlike public leaderboards that allow developers to train models specifically to pass known tests, Vals does not disclose its test materials. This approach is intended to prevent the gaming of metrics and provide enterprises with a clearer picture of how models might behave if deployed in real-world settings. The startup is also expanding its evaluation scope to include sensitive domains such as cybersecurity, biosecurity, and the application of international humanitarian law.

Scaling Operations and Industry Adoption

Vals has experienced rapid growth, with revenue increasing eightfold over the past year and its headcount expanding from eight to 25 employees. The company plans to further increase its staff and move into larger office space to accommodate its growing operations. In addition to its work with private enterprises, the startup has launched a program to provide model evaluations to federal agencies.
As AI models become increasingly integrated into the global economy, the demand for standardized, trustworthy evaluation has intensified. Krishnan suggests that as more AI companies move toward public offerings, independent benchmarks will become a central component of public filings and investment decision-making. The company's current trajectory reflects a broader industry shift toward treating reliable benchmarking as essential infrastructure, comparable to compute resources or data management platforms, as the sector moves from experimental phases to widespread production deployment.
The Indus Pulse is committed to accuracy and transparency.