I'm really not convinced that these models are even that much more intelligent, as opposed to simply being more token aggressive. I do not find Fable that much smarter than Opus 4.6, and no Opus model seems to have improved things much at all.
Benchmarks seem gamed at this point, real world experience just doesn't match up.
Benchmarks seem gamed at this point, real world experience just doesn't match up.