Shaw (spirit/acc)
Shaw (spirit/acc)|7月 29, 2026 18:13
RLHF is not reinforcement learning “Reinforcement training” is much more accurate to describe RLHF, GRPO, RLVR etc There is no learning going on You’re calculating reward instead of loss. Okay. But it’s not like the agent is learning from its own experience. It’s a statistical function over massive number of rollouts. If people feel really attached to RL for offline application of reward function to rollouts then I propose “experience learning” as the thing we all really mean— agents learning from their own experience and updating from it We can saturate LLMs up to AGI level but not far beyond. Reinforcement training is still limited to the tokens in distribution which came from human data, which means it’s a slot machine of exploring less likely tokens over time. We can never get superintelligence out of this approach, since novel tokens are not within the top k. Superintelligence requires an agent that can learn from its own experience. Maybe it can share experiences with other agents, but it’s critical that it can learn and update from its own experience or it has no capacity to explore beyond human-trained likely tokens. Tokens are too discrete to enable credit assignment in individual agents. Tokens can be the output of a continuous process (like when humans vocalize words) but they are an extremely inefficient structure for RL training. There is clearly another way— we are existence proof of this— the question is what the simplest model is that can embody all of the qualities never for an agent that can learn from its experience.(Shaw (spirit/acc))
+6
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads