And once again, Qwen 3.8 27B beats Opus 4.6, what the hell.
It's both funny and a bit terrifying and I still can't quite believe it. It runs decently on a gaming PC! Opus 4.6 came out only 6 months ago and was then broadly considered the new SOTA by a comfortable margin! How in hell did they package capability in the ballpark of a Feb 2026 frontier SOTA into 27B?!
More importantly, what's the point of building monster-scale data centers on unprecedented amounts of debt when a more than good enough model runs on a GPU from a couple years ago?
The coming months are going to be exciting, that's for sure...
I'm reminded of that paradox from sci-fi that says that starting an interstellar journey as soon as the technology is capable of it is uneconomical, because the trip will take so long that newer technology will arrive at the destination first, despite departing at a later date.
Very similar performance to 4.6 and codex 5.3 but slow and token inefficient. Still wildly impressive. Initially I didn't believe the results because 3.6 27b couldn't complete the benchmark so this is a massive leap in capability.
> More importantly, what's the point of building monster-scale data centers on unprecedented amounts of debt when a more than good enough model runs on a GPU from a couple years ago?
Probably because the future "monster" models will be insane. 100T+ param models might be the type of things that can independently run a small business, which means anyone not using them is at a distinct disadvantage to their competitors.
The top model from 2025 looks silly compared to the top model of the first half of 2026. Do you feel like progress has stalled?
> The top model from 2025 looks silly compared to the top model of the first half of 2026. Do you feel like progress has stalled?
I do. Pre-training is where the industry saw the “emergent properties” of LLMs arise and, for a time, people thought you could just keep scaling up bigger and bigger models but then the incremental gains from doing this did plateau. Labs will still do bigger models (Bytedance has a 10T planned), but these are sparse architectures and they aren’t going to have mind-blowingly greater intelligence. Fable didn’t either.
Scaling pre-training tokens also flattened out. What is still delivering gains is scaling RL on verifiable tasks. But that’s not general intelligence - it’s fitting models to specific tasks, which ML has always been good at. More importantly, most tasks to which humans apply their intelligence don't have computationally verifiable answers.
they didn't plateau, the hardware was effectively saturated and it wasn't until later in 2026 that newer hardware came online to provide enough capacity to keep scaling up the models efficiently
The gains from increased parameter scaling are sublinear: there's no more hockey-stick improvement to be seen going in that direction. That doesn't mean some improvement isn't possible - it's just going to be increasingly not worth doing.
Also, I think the fact that small open-weights models are catching up to the frontier rather than the frontier rapidly pulling away is evidence of this. In fact, by far the most dramatic capability increase story over the past two years has been the gains made in the small-parameter regime.
One might think, "hey, this agentic coding thing was a pretty big deal!", but I think it's a bit of a distraction because models only recently became optimized for this specific use case. It's not like they suddenly gained so much general intelligence that they magically had the ability to use a coding harness. No, the labs started spinning up a bunch of RL environments and generating rewards over long-horizon trajectories of combining these tools. It's an excellent application of LLMs but care needs to be taken interpreting how much "progress" has been mae in terms of raw generalized capability.
It's both funny and a bit terrifying and I still can't quite believe it. It runs decently on a gaming PC! Opus 4.6 came out only 6 months ago and was then broadly considered the new SOTA by a comfortable margin! How in hell did they package capability in the ballpark of a Feb 2026 frontier SOTA into 27B?!
More importantly, what's the point of building monster-scale data centers on unprecedented amounts of debt when a more than good enough model runs on a GPU from a couple years ago?
The coming months are going to be exciting, that's for sure...