Two names worth knowing
The paper this path follows is short, has no experiments, and is written by two people whose track record makes it hard to ignore.
Richard Sutton co-wrote the standard textbook on reinforcement learning and spent four decades arguing that general methods which scale with computation eventually beat methods that build in human knowledge. David Silver led the work behind AlphaGo and AlphaZero, the systems that beat the best human Go players and then, in AlphaZero's case, learned to do so from self-play without studying human games at all.
So when they publish an argument that the field has taken a wrong turn by relying on human data, the argument carries the weight of two people who have already demonstrated the alternative working in a narrow domain.
The paper is called Welcome to the Era of Experience. It is a position paper: a claim about where the field should go, offered as an argument rather than a result.
This path treats it as an argument to be examined. That means understanding it properly first, and then asking what would have to be true for it to be right.

