Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The benchmarks reported by Artificial Analysis are really weird in context of the ARC-AGI 3 scores and 'not not AGI' statements. It's an outright regression on the AA Agent composite vs GPT 5.6 Sol while a fraction of a point better on the full composite index. Could be the case it's just not showing up in benchmarks, for a good while Anthropic persistently trailed in benchmarks but had people swearing by it.
 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: