OpenAI Chief Scientist Jakub Pachocki: The Progress, Metrics, and Future of AGI

CN
10 hours ago

Written by: Techub News Compilation

Introduction

This conversation comes from the official OpenAI podcast, where host Andrew Main invited the company's Chief Scientist Jakub Pachocki and Research Scientist Simon Gdor. As core planners and deep participants in OpenAI's research roadmap, they analyze the fundamental questions in the field of artificial intelligence from a technical perspective: How do we define and measure progress in intelligence? How far are we from AGI (Artificial General Intelligence)? What do recent breakthroughs, such as models winning gold medals at the International Mathematical Olympiad (IMO), signify? This dialogue not only organizes the current state of technology but also deeply reflects on research and development philosophy, future breakthrough directions, and societal impacts.

Summary

  • The Measurement Dilemma: Traditional standardized tests are nearing "saturation," and models with performance matching or even surpassing humans necessitate new evaluation methods that reflect practical value and generalization ability.
  • Breakthrough Progress: Achievements like the IMO gold medal and programming competitions signify qualitative leaps in the model's capacity for complex reasoning and creative problem-solving, emphasizing "pure reasoning" rather than relying on tools.
  • AGI Pathways and Future: One of the core goals of AGI is "automated research," where AI autonomously discovers new knowledge and creates new technologies, greatly accelerating technological advancement but also presenting immense technical and societal challenges.
  • Perception of Development Speed: The progress felt within the industry far exceeds external perceptions; the absence of new records in the short term (like within a few weeks) does not mean a bottleneck has been reached; the long-term trend is exponential.
  • Advice for Learners: Programming (or more broadly, the ability to decompose complex problems) remains highly valuable in the future, and understanding system principles enables better mastery of AI tools.

From School to OpenAI: The Origin and Legacy of Interest

Jakub Pachocki and Simon Gdor's friendship began at the same school in Poland. They recall that a computer science teacher named Richard Ssrtovsky had a profound impact on them. This teacher had extensive experience as a programmer, particularly emphasizing programming competitions and the pursuit of excellence, introducing materials like graph theory and matrices far beyond the normal school curriculum.

Jakub Pachocki believes that tools like ChatGPT may allow more people to engage in similar in-depth explorations more easily. He cited his example of using ChatGPT to create an interactive version of the “Monty Hall problem” to aid understanding. This dynamic, multimedia explanation mode showcases AI's potential to surpass text explanations and create new types of learning materials. However, he also emphasizes that excellent teachers provide not just knowledge, but emotional support and ongoing interaction, aspects that current AI cannot fully replace.

Now, as the Chief Scientist at OpenAI, Jakub Pachocki primarily focuses on developing the company's research roadmap and determining the focal direction for technological pathways and long-term research planning. Meanwhile, Simon Gdor works on diverse cutting-edge exploratory tasks as a research scientist.

Defining and Measuring AGI: From Abstract Concepts to Concrete Milestones

When asked how to explain AGI to the average person, Jakub Pachocki did not provide a rigid definition but started discussing the evolution of capabilities. A few years ago, despite the broad prospects of deep learning, AGI seemed abstract and distant. Now, AI can engage in natural conversations on a wide range of topics, solve mathematical problems, and even conduct research. He realizes that human intelligence—communication, problem-solving, research—appears unified, but with technological advancement, we discover these are actually different capability modules.

Measuring progress is becoming increasingly difficult. Jakub Pachocki pointed out that the past scaling paradigm from GPT-1 to GPT-4 allowed standardized tests to effectively measure development levels. However, now, model performance on many tests is close to or has reached the human ceiling, exhibiting a "saturation" phenomenon. Additionally, specialized training on certain skills (like mathematics) can enable models to perform exceptionally well on related tests, but this does not equate to their general intelligence being equally strong in other areas. Therefore, judging a model's "quality" based solely on test scores has become ineffective; we need to focus more on its practical utility and problem-solving ability.

So, what constitutes an effective milestone? The IMO gold medal is viewed as an important benchmark. Jakub Pachocki explains that IMO problems do not require a vast amount of prior knowledge but severely test the ability to think deeply and solve problems creatively within one or two hours. Winning a gold medal means the model can purely reason (without using calculators or other tools) to solve highly creative mathematical problems. Particularly challenging is "Problem 6," which typically requires thinking outside the box; solving it is considered a more formidable achievement. Simon Gdor added an interesting detail: their model can "recognize" when it encounters unsolvable problems, demonstrating that it possesses self-awareness of its lack of progress, rather than forcing an incorrect answer (referred to as "hallucination"), indicating significant improvement in its judgment of its own cognitive boundaries.

Beyond mathematics, programming competitions are also key benchmarks. They mentioned a famous algorithm competition (Otkar) held in Japan that is open globally. In this competition, their model competed alongside top human players in a 10-hour event, ultimately securing second place. Simon Gdor shared an anecdote: the human player who won the competition remarked afterward, exhausted, that the AI model “really performed poorly,” as he had to compete fiercely against it until the last moment. These achievements in highly specialized, recognized tough fields provide concrete metrics for assessing AI's reasoning and problem-solving capabilities.

The Remarkable Pace of Advancement and the Vision of "Automated Research"

Simon Gdor expressed surprise that the outside world sometimes underestimates the speed of AI advancement. Recalling his personal experience, he noted that about 10 years ago, the natural language processing based on deep learning performed poorly, even misjudging sentiment in sentences like "The movie is not bad." From generating meaningful paragraphs with GPT-1 and GPT-2 to GPT-3 and GPT-4 becoming "personal AIs" that produce surprises, and now models participating in challenging programming competitions, the progress from an internal perspective is “astonishingly exponential.” He believes that if the progress made a decade ago could be quantified as 0.0X%, then discussing single-digit economic impact percentages now underestimates potential, with this number likely to rapidly rise to 10%, 20%, or even higher in the coming years. This is similar to the early days of the internet, where its vast impact was not immediately reflected in macroeconomic data.

Regarding OpenAI's research direction, Jakub Pachocki clearly stated that the company's goal is to create Artificial General Intelligence (AGI). One of their focal visions is “automated research.” He envisions a future where there exists a highly automated, AI-driven “researcher” or team capable of interacting with humans, absorbing information, running experiments, developing new technologies, writing code, and advancing projects, thereby “fundamentally accelerating technological progress.” This is not about applying AI to a specific field but leveraging its versatility to explore and create. He believes that while encouraging results have been seen in fields like medicine that require deep specialized knowledge and reasoning, the truly revolutionary discoveries and potential for advancements lie in “generality.”

Achieving this vision requires models to possess the ability to engage in prolonged, continuous thinking and computations on significant issues. Jakub Pachocki predicts that the next breakthrough direction may include significantly expanding the model's planning and reasoning “horizon,” allowing the model to dedicate vastly more computational resources and sustained thinking time to a critically important complex problem (like medical research or developing the next generation of models) than currently utilized. From GPT-4 to future models, the computational amount used for each response may experience orders of magnitude growth in exchange for higher-quality results.

Future Outlook: Interaction, Trust, and Advice for the New Generation

When discussing the experience of ordinary users interacting with AI in the coming years, Jakub Pachocki reiterated the macro impact that automated research will bring. Simon Gdor, however, focused more on the interaction interface itself; he believes that as AI becomes more robust, it will express itself in more diverse forms (text, voice, etc.), potentially enhancing the “sense of connection” between humans and AI, which will become an important social discussion topic.

Trust and safety are core challenges accompanying the growth in capability. Simon Gdor used his own allowance for the AI to read his Gmail calendar as an example, illustrating that user trust in models is increasing. However, he also pointedly noted that we have not yet reached a reliability level where we can fully trust models to ensure they are not misused. There exists a severe trade-off between enormous economic value and access to personal data, which requires ongoing efforts from professionals to address.

Finally, for today’s youth (especially high school students), both scientists offered sincere advice. Jakub Pachocki strongly recommended learning programming, or more accurately, cultivating the ability to “decompose complex problems into smaller parts.” He believes that even if AI can write code in the future, understanding how systems work and mastering the mindset of problem decomposition will give individuals an advantage in navigating AI. This is akin to a pilot needing to understand aerodynamics. He dismissed the argument of “no need to learn programming” as shortsighted and unreasonable.

Simon Gdor encouraged young people to “dare to dream and take action to change the world.” He shared his journey from Poland to Silicon Valley, finding that many self-imposed limitations do not actually exist. He has been greatly influenced by Paul Graham's “Hackers and Painters,” while Jakub Pachocki humorously admitted that the movie “Iron Man” sparked his interest in robotics, although he later found the reality differed greatly from the film and ultimately shifted to deep learning. Both believe that the emergence of AlphaGo was a turning point; it demonstrated the incredible potential of self-learning systems and continues to inspire their work.

The conversation ended on a light note but left profound reflections: we stand on the threshold where technological advancement could redefine “innovation” itself; how to measure, guide, and responsibly develop this power is a shared challenge facing the entire industry and society.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink