Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

That's not defensible.
 help



It absolutely is.

What has improved isn't the models, it's the harnesses.

Give GPT-3.5 a 1M context window and a modern harness, and you won't see any meaningful difference with Opus 5.

It's a bit hard to try with such old models, but for example I use Opus 5 / Fable at work and Sonnet 4.5 at home (because it's free via Amazon Q), and there's absolutely 0 difference in performance. None. Obviously 4.5 is only a year old, not 3, but try with any older model that has a decent context window and you'll get the same results.

In fact I'll go further than this and say that models are currently regressing. Opus 5 is much much worse than Opus 4.6 for example, and it's clear that Anthropic (at least - I don't use OpenAI models much) is just tokenmaxing rather than optimizing for performance.


Benchmarks are far from everything, but I would love to see the outcome of an experiment benchmarking GPT-4o (which is one of the earlier models with a >100k context window) against GPT-5.6 or Opus 5 in modern harnesses.

>Give GPT-3.5 a 1M context window and a modern harness, and you won't see any meaningful difference with Opus 5.

>models are currently regressing. Opus 5 is much much worse than Opus 4.6

I'm in sheer awe at these takes. Literally beyond parody.


Striking argument.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: