GLM 5.3 is probably the sweet spot open weights model if you want to go beyond deepseek flash or the new glm flash. I used it with pi and had a fairly good time, especially since it’s less touchy about cyber and whatnot than the US guys. It’s slightly behind Kimi in ability but it’s a lot easier to run it, I’d expect prices (and speed!) from third parties to be noticeably better.
Assuming you’re willing to drop a fat stack of cash on the upcoming Mac m5 ultra with 512 gb unified memory, you can even run it locally, quantized to 4 bit. Whether it’s even slightly reasonable, well, my wife would probably skin me alive but maybe yours is more understanding.
revolvingthrow
I'd like to ask Sam Altman if he still thinks that it's too dangerous to publish GPT-3. I mean, no one would use it, but what is his reasoning for not publishing it now, in 2026?
nkmnz
I previously posted that DS4Flash was _good_ but not _great_ on two DGX Sparks, but I have to say that GLM-5.3 is pretty amazing. It's been able to tackle all the random hard problems I've thrown at it and it has the intuition that DS4Flash seems to lack.
We're nowhere near a Fable-class model IMO, but things are going to get interesting in this next year.
mmastrac
What's very promising here is the number of tokens-vs-accuracy ratio. I am assuming their "output tokens" means tokens generated as part of thinking and any tool calls (what are referred to as "input tokens" from billing PoV by service providers). The Chinese models like Qwen3.8 and GLM 5.2 are insanely overthinking in our workloads (which are highly complex data analysis tasks). It's a factor of 3-4x over Opus and GPT models. Even with cheaper prices per 1M tokens, the cost ends up being higher, in some cases 2x. So this is very promising from GLM 5.3. Looking forward to trying it.
armcat
This might be the ideal form factor for local+ models. Even though I am a diehard qwen3.6-27B fan, it is not the same as this guy. GLM 5.3 you can actually run on-prem reasonably well, and it is a legit, proper, work horse. As in, this can do real work.
ThouYS
I've been using it more and more. Feels like Opus 4.8, in the best possible way.
scosman
I find GLM 5.3 Flash more interesting than 5.3. The fact 5.3 does not have vision is kind of a deal breaker. Also 5.3 Flash seems to be better at making pretty UIs.
The interesting part is that GLM-5.3 uses the same base model as 5.2, with the gains coming from post-training. It suggests better environments, verifiers and training trajectories may matter as much as another huge pretraining run.
Jeeetendra
I really like this model. It's the only one I've found so far (except maybe Kimi K3) that I would say I'm happy to use as a main driver.
comments (10)
Assuming you’re willing to drop a fat stack of cash on the upcoming Mac m5 ultra with 512 gb unified memory, you can even run it locally, quantized to 4 bit. Whether it’s even slightly reasonable, well, my wife would probably skin me alive but maybe yours is more understanding.
revolvingthrow
nkmnz
We're nowhere near a Fable-class model IMO, but things are going to get interesting in this next year.
mmastrac
armcat
ThouYS
scosman
redox99
fra
Jeeetendra
nullbio