Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

GPT usually performs better on DeepSWE while Claude does better on FrontierCode. These two coding benchmarks are pretty much the only ones right now that's still worth taking a look at imo.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: