a16z|Sep 09, 2026 14:48
Vals AI co-founder and CEO Rayan Krishnan with a16z's Ben Horowitz and Jennifer Li on grading AI, what it costs, and who gets to make the rules:
Every big industry eventually grows an independent testing layer. AI has credit ratings to learn from and Enron to avoid. Model capability today is still mostly self-reported.
As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations.
The harder problem is geopolitical. Reagan's "trust but verify" worked during the Cold War because you could fly over and count the missiles. No simple equivalent for AI models exists.
In this conversation with Erik Torenberg, they get into how you measure a model's ability to improve itself, why every good benchmark eventually has to be retired, and what happens when token spend begins to rival employee salaries.
00:00 Intro
02:20 Llama 4 on public vs private benchmarks
05:24 Nobody agreed how to test humans either
06:55 What movie ratings teach us about AI
08:55 The Enron problem in benchmarking
11:36 Why a good benchmark has to be retired
13:22 Evals that run for weeks, not seconds
16:20 Where the real workday starts at 4pm
18:08 A firm really is just its evals
20:35 Why Sonnet can cost more than Opus
22:42 One engineer, 6 billion tokens in a day
25:05 Who should set the rules for models
28:55 Public sector enforces, private verifies
33:32 Why sovereign AI is inefficient and happening anyway
35:00 The AI version of trust-but-verify
37:15 Where cyber evals have to go next
YouTube: https://www.youtube.com/watch?v=WO9c9qxDxzU
@RayanKrishnan @ValsAI @bhorowitz @JenniferHli @eriktorenberg(a16z)
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink