Sources
Burny — Effective OmniProgress on $1M ARC-AGI benchmark that is very hard for LLMs by carefully-crafted few-shot prompt to generate many possible Python programs to implement the transformations, generating ~5k guesses, selecting the best ones using the examples, and a debugging step, which is… https://t.co/jCfuY1fsps
Burny — Effective OmniProgress on $1M ARC-AGI benchmark that is very hard for LLMs by carefully-crafted few-shot prompt to generate many possible Python programs to implement the transformations, generating ~5k guesses, selecting the best ones using the examples, and a debugging step. https://t.co/jCfuY1fsps
Dileep George50% on ARC-AGI with GPT-4o This wonderful blog post brings out another point that I didn't explicitly mention in my blog -- ARC-AGI gets solved with a bunch of very clever tricks around existing models, and more search compute. https://t.co/YvoT4PC3yz https://t.co/CeXqixsbSF



