律动BlockBeats|Oct 10, 2026 10:03
[Yang Likun's Vision from 4 Years Ago Delivered: H-JEPA Maze Success Rate Rises from 18% to 73%]
Beating AI News Flash: AMI Labs, founded by Yang Likun, in collaboration with multiple universities, has unveiled H-JEPA, a world model capable of hierarchical future prediction and action planning. The higher level is responsible for long-term goals, while the lower level handles specific actions. The paper, code, and pre-trained weights have all been made public. H-JEPA builds upon the concept of JEPA. It first converts visual frames into abstract environmental features, then predicts future states based on actions, without the need to generate video frame by frame. However, long-term planning and immediate control require different types of information. For instance, when a quadruped robot navigates a maze, route selection primarily depends on position and pathways, while stepping requires attention to body posture. A single-layer model uses the same set of features for both tasks, making long-term planning increasingly challenging. H-JEPA combines multiple JEPAs into different planning hierarchies. The higher level predicts further into the future, passing intermediate goals to the lower levels, which then refine them into specific actions. Unlike previous HWM models, where all layers shared the same set of features, H-JEPA allows each hierarchy to learn its own environmental features. In maze experiments, the higher level retained positional information while gradually ignoring short-term details like leg posture. Yang Likun proposed a similar concept as early as 2022, and now the team has delivered an end-to-end trainable implementation. In the Visual AntMaze simulated maze, the three-level H-JEPA achieved a success rate of 73.3%, compared to 18.0% for the single-layer LeWorldModel, while requiring less computational effort for planning. However, more planning layers are not always better. In the Push-T task, which involves simulating object pushing, the three-layer model performed worse than the two-layer model. The team also conducted offline tests using the real robot video dataset DROID, where the predicted action trajectories were closer to expert demonstrations compared to the single-layer model. [Original Article Link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink