rick awsb ($people, $people)
rick awsb ($people, $people)|Sep 03, 2026 20:20
The renowned Vlasis Sholey, founder of arc agi, has just released the evaluation results of OpenAI GPT-6 Astra on the new benchmark ARC-AGI-3. Recorded a remarkable technological leap: Dual track testing performance: In a pure neutral, publicly only note taking standard testing environment (Standard Harness), Astra (max inference level) achieved an accuracy of 62.7%; In the Provider Adapter Harness environment that preserves the underlying state and compresses the context of open vendors, Astra (high inference level) achieved a score of 99.9%. Inference Steps and Action Efficiency: In terms of action efficiency, Astra demonstrates anti common sense efficiency - high inference levels complete levels with very few exploration steps, while overall API calls and token consumption are lower than mid-range inference. Compared to the median of 500 human subjects, Astra had lower step counts than humans in 96.0% of tasks, with an average reduction of 51.7% in action steps. The emergence of spontaneous tools and symbolic modeling: In the sandbox red team testing framework PRO-LONG, the model spontaneously constructs a compact domain specific algebraic language (DSL) shorthand when facing unknown rules, and independently writes scripts such as maze path solver (maze_solver. py) and state synchronization script (sync_date. py). However, Xiao Lei believes that ARC-AGI-3 is essentially a fully regulated, closed, and deterministic discrete toy environment; Even if AI hits this list, it doesn't necessarily mean that Astra is already AGI. Because human intelligence still maintains orders of magnitude (OOM) lag in learning time scales, world model complexity, and action space. But if we extrapolate linearly based on the speed of model development, from the perspective of the evolution path of computation and system architecture, these three dimensions do not have insurmountable thresholds (And the speed of model development is not linear): Action space with the lowest threshold: In the field of digital symbols, the combination of AI's instantaneous scheduling of massive APIs and code execution has already surpassed human freedom; In the field of physical embodiment, with the evolution of end-to-end diffusion strategies and VLA models, continuous multi degree of freedom control covering humans is mainly constrained by hardware and engineering data, and there are no theoretical blind spots. Learning time scales to compensate for individual capabilities through concurrency: AI can compress tens of thousands of years of interaction experience within weeks through cluster simulation and parallel self play, to compensate for the ability that AI currently does not possess, similar to the seamless accumulation of cognitive abilities in the human brain over decades. The complexity of the world model is the most difficult part: ARC-AGI-3 is a deterministic, discrete closed sandbox, while the real world is full of open randomness and irreversible physical laws. Relying on passive statistical fitting (Next token/frame) can easily accumulate long chain illusions; Humans rely on active physical interventions such as grasping and throwing to construct causality, while AI needs to bridge the complexity gap of 4-5 orders of magnitude by delving into physical reality to bear the real trial and error costs. This may require advances in robotics technology, requiring a large amount of physical world data and simulations. But ultimately, there are still no theoretical gaps, only engineering challenges. So no matter how experts define agi, we are either already in it or standing on the eve of agi's arrival
+1
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads