头雁
头雁|Oct 07, 2026 11:14
Another important company in the field of continuous learning is Richard Sutton (2024 Turing Award winner), who founded Oak Lab in Canada with former student Khurram Javed in July this year (both previously worked at John Carmack's Keen Technologies). For some unknown reason, he chose to leave Keen, whose direction is also RL+continuous learning for physical world models. The Oak Lab direction is based on reinforcement learning and continuous learning from real-time experience for intelligent agents, with one of its goals being "trillion parameters, real-time learning and planning, and approximately 20 watts of power consumption". To delve deeper into this company, one must look at some of Richard Sutton's previous research. OaK is a model-based RL architecture that emphasizes continuous learning and abstract discovery. There are three main characteristics: 1/Continuous learning of all components: There is no "freeze after training" phase, and the agent continuously updates from experience during runtime. 2/Each weight has a dedicated step size parameter, and meta learning is performed through online cross validation for better generalization. 3/Continuously create state and time abstractions, progress through FC-STOMP (five step loop): Feature Construction: Generate new features from existing state features. Placing a Subtask: Define sub problems based on high ranking features (such as "implementing a certain feature"). It is a termination condition. Learning an Option: Resolve the strategy and termination condition for this subtask. Option is a pair of (π, γ), where π is the policy (mapping of state to action distribution) and γ is the termination condition. Learning a Model of the Option: Predicting the consequences of executing the Option until termination (high-level transition model). Planning using the option's model: Based on these abstractions for higher-level planning. To understand some information about his system, you can study: Reward-Respecting Subtasks for Model-Based Reinforcement Learning OaK's core innovation: The key to OaK's continuous creation of abstractions (FC-STOMP loop) is that it constantly creates state abstractions and time abstractions on its own. This process is called FC-STOMP Progress (sometimes abbreviated as STOMP): Feature Construction 1/Create new state features from existing features. 2/Placing a Subtask For features with fluctuating weights, automatically propose a subtask to achieve respect for reward features. 3/Learning an Option Resolve subtasks to obtain an advanced skill (time abstraction) that can last for multiple time steps. 4/Learning a Model of the Option Predict the state and accumulated rewards after executing this option. 5/Planning using the Option's Model Use these advanced models for more efficient planning, and in turn improve strategies and value functions. Meta learning: It cannot be solved by manually adjusting the learning rate. The system must learn on its own when to learn quickly and when to learn slowly. Why does OaK emphasize meta learning in particular? It cannot be solved by manually adjusting the learning rate. The system must learn on its own when to learn quickly and when to learn slowly. Because intelligent agents require lifelong continuous learning: -The environment will change -New features and options will continue to emerge -Old knowledge sometimes needs to be quickly modified, and sometimes it needs to be protected It cannot be solved by manually adjusting the learning rate. The system must learn on its own when to learn quickly and when to learn slowly. After my research, I feel that: In short, there are still many differences in this company, and many concepts are not included in LLM, so I classified this company as an RL company. In addition, based on my observation, he wants to automate many aspects, such as generating subtasks and using the concept of Option to package a segment of behavior that spans multiple time steps into a whole: a continuous action. For example, in meta learning, without manually setting parameters, the core is that many subtasks are automated and cannot be manually adjusted for all dynamically generated subtasks. Including some of their latest backpropagation algorithms. The Alberta Plan for AI Research https://arxiv.org/abs/2208.11173 Reward-Respecting Subtasks for Model-Based Reinforcement Learning https://arxiv.org/abs/2202.03466 The Option-Critic Architecture https://arxiv.org/abs/1609.05140 OptionZero: Planning with Learned Options Basic Implementation of Options Framework https://(github.com)/theophilegervet/options-hierarchical-rl
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads