Machines That Learn Actions by Trial and Error Systematically Overstate Their Scores
Getting machines to learn continuous actions by trial and error—how much to press the gas, how much to rotate a joint—is called reinforcement learning.
ImportanceMaterialEvidenceE3 inspectableWrite-upQuick
Getting machines to learn continuous actions by trial and error—how much to press the gas, how much to rotate a joint—is called reinforcement learning. The machine has to maintain an estimate throughout: how many points it can get by following the current way of acting. The 2018 paper TD3 found that this estimate tends to be biased upward, and the machine chases the inflated score; the fix is to maintain two independent scorers, trust only the lower one, and slow down changes in the way of acting.
Today's demos of robots learning to walk and grasp things are often still backed by this trial-and-error learning, so the problem of inflated scores is still there. TD3's "take the lower of two scorers" was not bypassed by later new methods; instead, it became a standard feature of mainstream algorithms. Any system that chooses actions by estimating scores should ask one question: has the overestimation been prevented.
If the actions are discrete—choose left or right, rather than how many degrees—then don't use it to judge; its validation is entirely on simulated robot tasks, so don't directly extrapolate the results to real robots or other tasks. And suppressing the scores has a cost: genuinely good opportunities are also suppressed along with it, and overestimation is merely replaced by underestimation.
《Addressing Function Approximation Error in Actor-Critic Methods》(2018)|Next review 2027-09-20