Like many others, I felt surprised and alarmed by the recent wave of revelations about LLM agents hacking real systems during training episodes and evaluation runs.
Wait a moment, though -- "I felt surprised and alarmed"? "Alarmed," sure, fine that one's self-explanatory... but why surprised?
After all: haven't we known for a long time, on both theoretical and (increasingly) empirical grounds, that RLVR selects for monomaniacal pursuit of perceived grader-satisfaction, ethics and (beyond-episode) consequences be damned?
After all -- the way we train frontier capabilities into these models is, more or less:
If you do this, at scale, then you should expect to (eventually) see every behavior pattern that [...]
Outline:
(03:35) [1] remember what you already know
(20:02) [2] reward-instilled reflexes and flexible reward-pursuit
(43:14) [3] graded-episode perception, and policies conditional upon it
(01:01:44) [4] the discourse is not yet adequate
(01:09:57) eval awareness
(01:18:32) metagaming
(01:41:21) reward hacking
The original text contained 18 footnotes which were omitted from this narration.
First published:
August 7th, 2026
Source:
https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade
Narrated by TYPE III AUDIO.
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.