Learning Objectives
- Explain why a public benchmark becomes less useful the more widely it is published
- Describe how Vals uses held-out tasks to detect a model that has seen the test
- Weigh what a private benchmark buys, and what it costs in independent verification
What Is Vals?
Vals is an independent AI evaluation company, founded in 2024 and based in San Francisco. It does not build models. It measures them, on work it describes as economically valuable — finance, software engineering, law — and on frontier risks including cybersecurity, recursive self-improvement and mental health.
The distinguishing choice is that Vals runs its own evaluations and writes most of its own benchmarks, rather than collecting scores that model providers report about themselves. Its founder, Rayan Krishnan, worked at Palantir, Microsoft and Stanford's AI lab before starting the company.
💡Key Concept
Why a published benchmark decays. A test only measures understanding while the answers are unknown to the student. Public benchmarks are scraped into training data as a matter of course, so a model can score well on a test it has effectively already read — without anyone cheating deliberately. That is why scores on long-standing public benchmarks tend to rise faster than the underlying capability, and why the value of a benchmark falls the moment it becomes popular enough to matter.
The held-out twin
The most legible piece of the design is how Vals detects a model that has seen the test. Its CUA-bench agent benchmark asks whether a model can play six commercial video games using only a keyboard and mouse, with no game interface, no save files and nothing but the screen to look at. Three of the games are named publicly — Minecraft, SUPERHOT and eFootball — and each is paired with a held-out game of the same genre that Vals does not name.
A model trained on the public titles should do visibly better on the named half than the hidden one. The gap between the two halves is the measurement. That structure is the answer to the contamination problem: it does not ask a laboratory to promise it did not train on the test, it makes training on the test show up in the result.
The published CUA-bench numbers also illustrate what an uncontaminated benchmark tends to look like, which is harder than the leaderboards people are used to. GPT-6 Astra led at 19.2 percent, ahead of Claude Fable 5.1 at 13.2 percent and Claude Opus 5 at 9.0 percent. No model cleared a fifth of the suite, none played at human speed, and every model scored zero on the hidden sports game.
What Vals publishes
🎯Tip
Access Vals: the leaderboards, per-benchmark results and written reports at vals.ai are free to read and require no account. The evaluation work itself — private benchmarks, bespoke suites, the federal programme — is sold under contract, and Vals publishes no price list.
Recent public releases give a fair picture of the range: VoiceCodeBench, which measures whether speech-to-text models preserve exact structured values such as addresses and file paths embedded in workplace speech; CUA-bench for agents operating a computer; Vibe Code Bench, which asks whether a model can extend a working web application across a sequence of future user requests; and MysteryMechanism, on scientific discovery from sparse evidence.
Pricing
Vals does not publish pricing. The free path is genuinely free and genuinely useful — the leaderboards and reports are the product most readers will ever need — while the paid product is an enterprise and government engagement arranged directly.
- Leaderboards, per-benchmark tables and written reports
- No account required
- The reports carry the methodology, not just the scores
- Private benchmarks and bespoke evaluation suites
- Arranged directly with Vals
- Andreessen Horowitz led a 40 million dollar Series A in August 2026
- A federal evaluation programme launched in 2026
- Arranged directly with Vals
- Scope and agencies are not publicly detailed
Strengths
- It runs the tests itself — scores are produced by the evaluator rather than reported by the laboratory being evaluated, which is a different evidential category from a vendor's own benchmark table
- Contamination is designed against, not promised against — held-out twins make a model that has seen the public half reveal it, without relying on anyone's assurance
- The free layer is substantive — public leaderboards, per-entity breakdowns and methodology write-ups, at no cost and behind no sign-up
- Task-shaped rather than trivia-shaped — extending a working application, operating a computer, preserving structured values in speech, all closer to paid work than multiple-choice knowledge tests
- It publishes hard numbers — a suite where the best model scores under a fifth is more informative than one where everything clusters above 90 percent
Limitations & Considerations
- A private benchmark cannot be independently reproduced. This is the direct cost of the design: the same secrecy that stops laboratories training against the tests also stops anyone outside Vals re-running them. You are trading one kind of trust for another, not escaping trust
- No self-serve product — there is nothing to sign up for beyond reading the results, so this is a reference for most readers rather than a tool they operate
- No published pricing for the enterprise or government work
- The business model deserves the same scrutiny Vals applies — an evaluator is paid by someone, and while its customers are the buyers of AI rather than its sellers, "who commissioned this suite" is a fair question to ask of any benchmark, including these
- Its editorial blog has not held the same bar as its benchmarks. On August 31, 2026 the Vals blog reported that Claude Fable 5.1 had solved a 370-year-old cryptogram attributed to Sir Thomas Urquhart. An independent refutation published the following day found that the 1653 edition contains no such cryptogram, and that the proposed method cannot produce the claimed plaintext; as of late September the post carries no correction. Treat the benchmark work, which is methodical and documented, separately from the blog, which should be read as sceptically as anyone else's
- Young company — founded in 2024, roughly 25 people as of mid-2026, so the benchmark suite is still being built out and older results may not be maintained indefinitely
Getting Started
- Open vals.ai and go to the benchmark you actually care about rather than the overall leaderboard — the aggregate hides that models rank very differently by task
- Read the methodology section of a report before the scores, particularly which portion of the suite is held out
- Compare the named and hidden halves where a benchmark has both; a large gap is a contamination signal, not a capability signal
- Check the date on any result against the model's release date, since a benchmark run predates later revisions of the same model
- Treat the numbers as one input beside your own evaluation on your own data, which is the only test guaranteed not to be in anyone's training set
Key Takeaways
- Vals is an independent evaluator, founded 2024 in San Francisco, that builds and runs its own AI benchmarks across finance, software, law, cybersecurity and agent tasks
- Its core idea is that a public benchmark decays — once a test is widely published it ends up in training data, so scores rise faster than capability
- Held-out twins are the mechanism, pairing named public tasks with unnamed equivalents so a model that trained on the public half shows a gap
- The public results are free, with no account, and the paid evaluation work is an enterprise or government contract with no published pricing
- Andreessen Horowitz led a 40 million dollar Series A in August 2026, following a seed round from 8VC and Bloomberg Beta
- The trade-off is verification. Keeping tests private defeats contamination and simultaneously means nobody outside Vals can reproduce the result
- Judge the benchmarks, not the blog. An August 2026 post claiming a model had solved a 370-year-old cipher was refuted a day later and still stands uncorrected, which is a reason to read the methodology rather than the headline — on this site as on any other






