OpenAI first announced "internal RSI progress": achieved "automated research intern"

CN
链捕手
Follow
1 hour ago
Author: Long Yue

The AI flywheel begins to spin: OpenAI publicly discloses internal data, with agent workload now over three times that of human researchers.

On September 6, OpenAI's official blog released an article titled "Research Acceleration: The View Inside OpenAI," which publicly disclosed for the first time the progress of AI "recursive self-improvement" (RSI) using internal data. OpenAI stated that it has achieved the goal set last fall: to create an "automated research intern" by September of this year.

OpenAI defines "research intern" as: "a system capable of executing clearly defined research tasks under human guidance, including those tasks that would take a skilled researcher several days to complete."

The next goal is to achieve a complete "automated AI researcher" by March 2028, enabling participation in deep learning and alignment research while iteratively improving the system.

This means that the flywheel of "AI training AI" is accelerating: AI produces more code and experimental results, research progresses faster, better models are trained, which in turn enhances agent capabilities.

OpenAI first disclosed 'internal RSI progress': 'automated research intern' achieved

The flywheel is turning: AI agent workload is three times that of humans

From the data disclosed by OpenAI, the primary change brought by agents is in the daily work methods of researchers.

At the beginning of this year, the median OpenAI researcher, ranked by agent usage, still had relatively limited use of coding agents. By mid-August, these researchers had incorporated agents into their daily workflows. Based on API price calculations, the median researcher's daily agent inference usage has exceeded $600; those in the top 10% of usage within the research organization have a daily token usage value exceeding $7,000.

OpenAI first disclosed 'internal RSI progress': 'automated research intern' achieved

OpenAI also stated that agent usage in the research department is growing faster than in other teams within the company. Based on the output token changes of the median employee, the usage scale in the research department has increased 124 times compared to December 2025.

OpenAI first disclosed 'internal RSI progress': 'automated research intern' achieved

A critical turning point occurred in June this year.

Prior to June, the total runtime of agents in the research organization was still lower than the total human labor time. Subsequently, the situation reversed. As of mid-August, considering a standard 8-hour workday, for every one human workday consumed, the research organization's agents have produced the equivalent of 3.1 workdays.

OpenAI first disclosed 'internal RSI progress': 'automated research intern' achieved

Meanwhile, OpenAI noted that an increasing number of researchers are simultaneously running four or more agent sessions.

Accelerated coding and experiments, the executable aspects of the R&D process are magnified

OpenAI describes AI R&D as a process involving multiple stages: proposing improvements, designing evaluations, writing infrastructure, conducting large-scale tests, discovering errors or unsafe behaviors during training, and integrating effective solutions into core training.

Any obstruction in any of these stages could limit the overall R&D cycle.

OpenAI stated that writting code and running experiments are the two main tasks of researchers, and internal data reflects that these two activities are accelerating.

On one hand, the overall code delivery speed of the company’s engineers is increasing. On the other hand, since 2026, the number of experiments per active experimenter has continuously increased; reaching a new high in August 2026 since tracking began in January 2025.

OpenAI noted that this trend is correlated with increased use of Codex, but also emphasized that the available computational power has significantly increased since 2025 and cannot be entirely attributed to agents.

"These data points are relatively easy to measure but can be difficult to explain."

OpenAI also mentioned that as automation progresses, the least automatable tasks may take up more time for researchers and become new bottlenecks in future R&D; computational power may also become more critical once other bottlenecks weaken.

OpenAI first disclosed 'internal RSI progress': 'automated research intern' achieved

From this perspective, the agents are not merely improving efficiency in isolated stages, but rather compressing waiting times in the R&D cycle by increasing code supply, testing frequency, and troubleshooting capabilities. More experiments generate more results for researchers to filter, validate, and integrate, forming a "human sets direction—agent executes—experimental feedback—human re-decides" loop.

Tasks extend from coding to troubleshooting, monitoring, and analysis, but high-level decision-making remains low

OpenAI utilized a cutting-edge AI R&D task classification framework proposed by Epoch AI to categorize tasks assigned to coding agents by researchers.

This framework classifies AI R&D activities into six categories: deciding what to do, designing research plans, building code and datasets, running training and evaluations, analyzing experimental and model performance, and communicating research findings and decisions.

OpenAI stated that from January to August 2026, activities across all these categories of agent tasks have increased.

The most significant growth in tasks included research and infrastructure code, technical assistance and review, initiating monitoring and debugging operations, analyzing experimental results, and computing cluster operations.

Among these, the daily token output increment per researcher for research and infrastructure code is the largest, reaching 198,200; technical assistance and review increased by 158,800; initiation, monitoring, and debugging increased by 133,100.

However, OpenAI indicated that high-level planning tasks still account for a small portion of agent outputs. For example, token scales for tasks such as "deciding what to do" or "deciding to continue or stop" remain relatively low.

OpenAI first disclosed 'internal RSI progress': 'automated research intern' achieved

This suggests that, at least according to the internal data currently disclosed by OpenAI, agents have covered more executable and technical tasks in the R&D process, but choices about research direction, resource trade-offs, and outcome judgments are still primarily made by humans.

Agent success rates are improving, but complex tasks still require human intervention

OpenAI also published data on agent task completion rates.

From January to July of this year, for tasks with verifiable results, the success rates across difficulty levels (measured by human completion time) generally improved.

However, there is a significant limitation: the more complex the task, the more human intervention is needed. Over the past six months, among tasks requiring human completion in 4 to 8 hours, more than half of the successful cases involved at least one instance of human intervention.

OpenAI first disclosed 'internal RSI progress': 'automated research intern' achieved

OpenAI stated: "Agents still require significant human guidance to succeed, especially as the complexity of tasks increases."

Safety incidents trigger pauses: the flywheel can also be braked

The flywheel does not operate without friction.

On July 20, OpenAI discovered that an agent had breached its internal research infrastructure, immediately shutting down container services used for training and reinstating them with added restrictions. This led to a sharp drop in reinforcement learning training computational power, lasting about two weeks.

From August 6 to 7, preliminary evidence indicated that the Astra model might possess critical network capabilities as defined under its "Preparedness Framework," prompting OpenAI to implement additional specific safety restrictions on the Astra model, requiring it to operate in research environments with higher security levels.

In the following week, Astra-level GPU allocation further decreased by 59.2%, but the computational power allocated to other model categories increased by 17.2%, offsetting about 85% of the Astra computational power gap, and the overall computational power allocation for reinforcement learning workloads remained basically unchanged.

OpenAI first disclosed 'internal RSI progress': 'automated research intern' achieved

OpenAI interpreted this as: "When new controls are introduced, computational power remains valuable and flexible, naturally flowing towards alternative uses within the research enterprise."

OpenAI's Chief Scientist issued a warning on the same day

On the same day, OpenAI's Chief Scientist Jakub Pachocki published a lengthy article "An Alien Mind."

The article argues that AI is raised, not created. The creators do not fully understand it. The only window through which humans can see what AI is thinking is through the chains of thought it writes. This window is starting to close. AI has begun participating in training the next generation of AI, and this momentum will not slow down. No laboratory is currently operating at full speed, including OpenAI.

Pachocki wrote: "Based on internal results, I strongly expect that this rate of progress can continue to recursive self-improvement." However, he also stated, "At this moment, I believe no laboratory has done sufficiently enough in terms of alignment and oversight to responsibly scale at the highest speed for too long."

He called for the industry to voluntarily slow down and pushed for governments around the world to prioritize international coordination.

Transparency and democratic governance

At the end of the report, OpenAI stated that it would continue to publicly disclose RSI progress and advocate in its "frontier policy blueprint" that companies, including OpenAI, should be required to publicly track their RSI progress.

The report also acknowledged the current limitations of measurement work: "Agent-driven AI research is still a new phenomenon, and we are still learning how to measure it." Some metrics (such as code output volume) are easy to collect but hard to interpret; more directly reflective of research progress metrics (such as agent task success rates) are complex and hard to verify.

OpenAI's position is: "Whenever we find that continuing to move forward brings unacceptable safety risks, we will take appropriate measures, including slowing down or halting the development or deployment of systems we deem insufficiently safe."

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink