
Magic AI Labs has introduced a new evaluation method called HashHop, designed to address the limitations of existing long-context evaluations. The HashHop algorithm allows for a context window of 100 million tokens while utilizing a fraction of the memory of a single H100 GPU. This innovative approach has garnered attention for its potential to enhance the evaluation of large context models. The team at Magic AI Labs has received praise for their work, which includes a blog post detailing the weaknesses of popular long-context evaluation methods and showcasing HashHop as a viable alternative. The new method is anticipated to replace outdated benchmarks, such as the 'Needle In A Haystack' test, with more effective standards for assessing long context windows.

