Sources
Rafael Rafailov @ NeurIPSAs I said earlier, we need to figure out RL at foundation model scale. This work is yet another piece of the missing puzzle. What I still wonder is how dynamic programming RL training affects the knowledge inherent within a pre-trained model? Some thoughts on this soon. https://t.co/qRDoufti5x
Ted XiaoReplacing regression with classification in RL improved performance across many domains, offline and online settings, and even scales to generalist settings like robotics! https://t.co/bRU2u2u82E
Jesse FarebrotherFraming regression as a classification has been “dark knowledge” for some time. We wanted to shed some light on this phenomenon in deep RL: Framing value-learning as a classification significantly improves performance and scalability in deep RL. But... not all classification… https://t.co/A0vbOexNpq

