Vals logo

Vals

Independent AI evaluation company that builds and runs its own benchmarks on economically valuable work — finance, software, law — and on frontier risks including cybersecurity and recursive self-improvement. Keeps test material private so vendors cannot train against it.

Share

Listen to this lesson

Free preview · first 0:30
0:00 / 0:30

Unlock audio and more

Audio streaming, downloadable PDFs and certificates come with Plus and Pro.

📋About Vals

Updated September 20, 2026

Vals is an independent AI evaluation company founded in 2024 and based in San Francisco. It builds and runs its own benchmarks rather than reporting scores supplied by model vendors, covering economically valuable work such as finance, software engineering and law alongside frontier risks including cybersecurity, recursive self-improvement and mental health. Its founder, Rayan Krishnan, worked at Palantir, Microsoft and Stanford's AI lab before starting the company.

The commercial argument is that a benchmark a laboratory can read is not a measurement. Public test sets leak into training data, whether deliberately or through ordinary web scraping, so a model can score well on a benchmark it has effectively already seen. Vals keeps its test material private and pairs public tasks with held-out equivalents — its CUA-bench agent benchmark, for instance, pairs three named commercial video games with unnamed games of the same genre, so a model trained on the public titles reveals it as a gap between the two halves. Results are published as public leaderboards and written reports that anyone can read, while the evaluation work itself is sold to enterprises and government.

Andreessen Horowitz led a $40 million Series A in August 2026, after a seed round from 8VC and Bloomberg Beta. The company has said revenue is roughly eight times what it was a year earlier and that headcount grew from 8 at the start of 2026 to about 25 by mid-year, with a federal evaluation programme among its more recent work.

Vals sits in a category that matters more as model claims get harder to check independently. Its closest peer in the catalogue is Arena, which crowdsources human preference votes into a public leaderboard, and Patronus AI, which sells evaluation and agent-testing infrastructure to teams building on models. Vals differs from both in running the evaluations itself and withholding the tests.

🛠️Products & Tools (1)

ValsAI Evaluation & Benchmarking

Independent AI benchmarks run in-house on finance, software, law, cybersecurity and agent tasks, with test material kept private so vendors cannot train against it. Public leaderboards and written reports are free to read; the evaluation work itself is sold to enterprises and government.

Keep track of the companies you’re watching

  • The AI Hub on a phone: a 12-day AI Skill Streak and an expanded Content updates alert listing the saved items that changed.
  • Recommended for you on a phone: nine personalised suggestions labelled Trending in AI news, On your saved list, and Popular.
  • My AI Tools on a phone: saved tools including GitHub Copilot and OpenAI Codex, each with an Updated badge.

Swipe for Recommended for you and My AI Tools

Your AI Hub — sample data.

📰Vals in the News

Showing the only story where Vals is tagged in Top AI Stories.