头雁|Oct 06, 2026 07:25
Podcast interview with the founder of Reflection @ reflection.ai, discussing his desire to learn from our Liang Sheng and build a Western DeepSeek
Alex Heath's Sources podcast interview (approximately 62 minutes): Guests are Reflection's two co founders: CEO Misha Laskin (formerly responsible for reward modeling for Google Gemini) and CTO Ioannis Antonoglou (involved in AlphaGo, AlphaZero, and leading Gemini's RLHF). The core is the company's upcoming release of its first open-source weight model, Beam, and their desire to create the 'DeepSeek of the West'.
Why open source models suddenly become popular:
They use real estate as a metaphor: closed source models are like renting a house, and open source weights are considered "possessing" intelligence. The AI commercial market has only been around for about three years, and in the long run, the West only has leasing options; As enterprises and application companies grow, ownership, controllability, and customizability become more important. They mentioned that on platforms such as OpenRouter, the proportion of open source tokens has increased from about 30% to about 70%. In addition, geopolitical factors between China and the United States are also pushing this matter forward.
The origin story of Reflection: The two worked together at DeepMind, with Ioannis being Misha's supervisor and working together on Gemini 1/1.5. Ioannis was responsible for RLHF. The company initially had two bets: RL could make coding and agent capabilities; The West will have a strong open-source foundation to continue conducting research. The previous bet was correct, but the latter failed after DeepSeek V3. After waiting for several months without an equivalent Western open-source model, I decided to create my own model instead. The origin is an awkward comedy dinner between Cellar and Lobster Roll during the New York sprint, where Ioannis first proposed to start a business together. Nvidia is an important investor and also supports open source, but they came to the conclusion of going open source themselves, which was not pushed by Jensen. Financing exceeded the $2 billion level, with a known valuation of approximately $25 billion at the time.
Why China leads in open source:
They believe that in the early days of the West, the closed source rental model developed first, driven by commerce; After Llama 3, there is a lack of commercial incentives to continue doing top-level open source. Chinese laboratories are more driven by geopolitical factors, and DeepSeek V3 is a turning point, pushing China's open-source intelligence to a globally relevant position. Now that the market has matured, these laboratories are also beginning to generate revenue.
What is Beam:
A MoE model with approximately 500 billion total parameters and 23 billion activation parameters, pre trained from scratch, not distilled. Positioning involves encoding, inference, and agent, benchmarking the strongest open source models on public benchmarks, emphasizing inference token efficiency (they claim that comparable open source cutting-edge models and GPT like models are several times more efficient). The target users are enterprises, public sectors, and sovereign clients, serving as a self hosted 'main model'. Technologically, they have increased the computational power of reinforcement learning to a high level (mentioning about 10000 GB300 running for several weeks), believing that this is the key to squeezing out intelligence from smaller models.
Training from scratch:
Not distillation. There are two reasons: to create a cutting-edge open source model that is truly built from scratch by Western teams; For large-scale RL to be stable, it must be self controlled through pre training. Pre training data needs to first cultivate reasoning ability, and MoE's routing and experts still need to maintain a balance under high computing power RL, so end-to-end self training is necessary to fit enough intelligence into a smaller scale. Last November, a pre training team was formed from scratch with only 5 people.
Security and openness:
They advocate openness for greater security, analogous to Linux, encryption protocols, and Linus's Law: vulnerabilities require a large amount of external scrutiny to be identified. There are too few safety researchers in the closed source laboratory, similar to 'insufficient white blood cells'. They also acknowledge that when their capabilities reach a certain threshold, both open source and closed source should be regulated and restricted from being released; Hugging Face turned to open source models for defense after being attacked, and they used it as an example of ecological complementarity. The alignment method is similar to that of a closed source laboratory: evaluation SFT、RL, Balance between usefulness and rejection.
Company and Business Model:
The two started their business after collaborating on Gemini at DeepMind, originally betting on reinforcement learning and having a Western open-source foundation; After DeepSeek, the Western open-source frontier retreated, and they switched to making models from scratch themselves. Financing of approximately $4.6 billion (Nvidia, Sequoia, Lightspeed, etc.); Valuation and computing power agreements were also mentioned in the interview, including Nebius and SpaceX.
The logic of making money is not selling weight itself, but using open-source models to drive reasoning and meet the needs of enterprises: helping companies, governments, and allies deploy reasoning agent、 Cluster and security tools, be a full stack partner with your own intelligence. Long term vertical integration of computing power may increase profits; Income is roughly equal to intelligence density x controllable computing power x customer trust.
Collaboration and Next Steps:
Mentioning about 250 MW data centers in South Korea and collaborating with infrastructure partners such as Dell: the other party provides land, electricity, and bare metal, while Reflection provides models and software stacks on them. Distribution is not just about throwing Hugging Face, but also about supporting open source tools and integrating with cloud and inference providers. In 2027, we plan to develop significantly larger models and conduct multiple rounds of RL iterations on the same base as other laboratories, with the goal of closing the gap between open source and closed source frontiers.
What should Reflection do next:
Our mission is to bridge the gap between open source and closed source frontiers. In 2027, significantly larger models will be released and RL will continue to increase; multiple rounds of releases will be made on the same base, similar to others' 5.1 and 5.2. The next generation is already doing it. They consider this RL as the largest one in open source, running around 10000 GB300 for several weeks without stopping learning, with the bottleneck being computing power. Scientific breakthroughs have been made, and the next step is to make scaling more efficient and distribute the benefits through multipolarity, transparency, and builder ecology.
https://sources.news/p/reflection-founders-open-weight-beam-release
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink