Sources
Loading...
Additional media
Loading...

Researchers from Google DeepMind, University of Toronto, MILA, and UCLA have introduced a novel approach called Generative Reward Modeling (GenRM). DeepMind's GenRM improves the accuracy of Large Language Models (LLMs) by training them to verify their own outputs using next-token prediction and chain-of-thought (CoT) reasoning. The approach leverages the text generation capabilities of LLMs to improve their performance.

