a16z
a16z|Aug 06, 2026 18:17
vLLM runs on half a million GPUs at any given moment. Most people have never heard of it. Simon Mo, co-founder and CEO of @inferact and lead maintainer of vLLM, sits down with a16z’s Matt Bornstein and Elena Burger to discuss what it takes to actually run open models in production, the advantages of open models, Simon’s mission at Inferact, and more. 00:00 Intro 01:46 When open source became critical infrastructure 08:55 Day zero model releases, and the drama behind them 14:59 What Kimi K3 actually buys you 18:56 Why open model licenses are changing 22:24 The pharmaceutical analogy for funding model training 26:16 If GPUs got 99% cheaper 29:08 Why vLLM, OpenRouter, and Ollama all started before ChatGPT 35:42 Building a company on an open source project 39:48 The inventor of RoPE removing RoPE @simon_mo_ @BornsteinMatt @VirtualElena(a16z)
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads