Agreed, these models seem relatively mediocre to Qwen3 / GLM 4.5

modeless · 2025-08-05T17:44:26 1754415866

Nah, these are much smaller models than Qwen3 and GLM 4.5 with similar performance. Fewer parameters and fewer bits per parameter. They are much more impressive and will run on garden variety gaming PCs at more than usable speed. I can't wait to try on my 4090 at home.

There's basically no reason to run other open source models now that these are available, at least for non-multimodal tasks.

tedivm · 2025-08-05T17:55:59 1754416559

Qwen3 has multiple variants ranging from larger (230B) than these models to significantly smaller (0.6b), with a huge number of options in between. For each of those models they also release quantized versions (your "fewer bits per parameter).

I'm still withholding judgement until I see benchmarks, but every point you tried to make regarding model size and parameter size is wrong. Qwen has more variety on every level, and performs extremely well. That's before getting into the MoE variants of the models.

modeless · 2025-08-05T18:04:49 1754417089

The benchmarks of the OpenAI models are comparable to the largest variants of other open models. The smaller variants of other open models are much worse.

mrbungie · 2025-08-05T18:13:01 1754417581

I would wait for neutral benchmarks before making any conclusions.

bigyabai · 2025-08-05T19:25:05 1754421905

With all due respect, you need to actually test out Qwen3 2507 or GLM 4.5 before making these sorts of claims. Both of them are comparable to OpenAI's largest models and even bench favorably to Deepseek and Opus: https://cdn-uploads.huggingface.co/production/uploads/62430a...

It's cool to see OpenAI throw their hat in the ring, but you're smoking straight hopium if you think there's "no reason to run other open source models now" in earnest. If OpenAI never released these models, the state-of-the-art would not look significantly different for local LLMs. This is almost a nothingburger if not for the simple novelty of OpenAI releasing an Open AI for once in their life.

modeless · 2025-08-05T19:46:54 1754423214

> Both of them are comparable to OpenAI's largest models and even bench favorably to Deepseek and Opus

So are/do the new OpenAI models, except they're much smaller.

UrineSqueegee · 2025-08-05T20:12:09 1754424729

I'd really wait for additional neutral benchmarks, I asked the 20b model on low reasoning effort which number is larger 9.9 or 9.11 and it got it wrong.

Qwen-0.6b gets it right.

bigyabai · 2025-08-05T22:36:21 1754433381

According to the early benchmarks, it's looking like you're just flat-out wrong: https://blog.brokk.ai/a-first-look-at-gpt-oss-120bs-coding-a...

nialv7 · 2025-08-07T02:28:17 1754533697

Looks OpenAI's first mover advantages are still alive and well

thegeomaster · 2025-08-05T20:13:23 1754424803

They have worse scores than recent open source releases on a number of agentic and coding benchmarks, so if absolute quality is what you're after and not just cost/efficiency, you'd probably still be running those models.

Let's not forget, this is a thinking model that has a significantly worse scores on Aider-Polyglot than the non-thinking Qwen3-235B-A22B-Instruct-2507, a worse TAUBench score than the smaller GLM-4.5 Air, and a worse SWE-Bench verified score than the (3x the size) GLM-4.5. So the results, at least in terms of benchmarks, are not really clear-cut.

From a vibes perspective, the non-reasoners Kimi-K2-Instruct and the aforementioned non-thinking Qwen3 235B are much better at frontend design. (Tested privately, but fully expecting DesignArena to back me up in the following weeks.)

OpenAI has delivered something astonishing for the size, for sure. But your claim is just an exaggeration. And OpenAI have, unsurprisingly, highlighted only the benchmarks where they do _really_ well.

sourcecodeplz · 2025-08-05T19:23:09 1754421789

From my initial web developer test on https://www.gpt-oss.com/ the 120b is kind of meh. Even qwen3-coder 30b-a3b is better. have to test more.

moralestapia · 2025-08-05T17:53:48 1754416428

You can always get your $0 back.

Imustaskforhelp · 2025-08-05T17:58:14 1754416694

I have never agreed with a comment so much but we are all addicted to open source models now.

recursive · 2025-08-05T20:08:42 1754424522

Not all of us. I've yet to get much use out of any of the models. This may be a personal failing. But still.

satvikpendem · 2025-08-05T18:04:14 1754417054

Depends on how much you paid for the hardware to run em on

cvadict · 2025-08-05T19:53:15 1754423595

Yes, but they are suuuuper safe. /s

So far I have mixed impressions, but they do indeed seem noticeably weaker than comparably-sized Qwen3 / GLM4.5 models. Part of the reason may be that the oai models do appear to be much more lobotomized than their Chinese counterparts (which are surprisingly uncensored). There's research showing that "aligning" a model makes it dumber.

xwolfi · 2025-08-06T02:42:09 1754448129

The censorship here in China is only about public discussions / spaces. You cannot like have a website telling you about the crimes of the party. But downloading some compressed matrix re-spouting the said crimes, nobody gives a damn.

We seem to censor organized large scale complaints and viral mind virii, but we never quite forbid people at home to read some generated knowledge from an obscure hard to use software.