Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I didn't accept a single edit from this model over the entire week, just saying. I do not understand how it's being benchmarked on par with Sol and other larger models.


It was certainly almost RL-fried to overfit the benchmarks, at the expense of actual usability. See Opus 5.


is it a benchmarkmaxxing model?!




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: