GLM-5.3-Flash

(z.ai)

200 points | by Philpax an hour ago ago

66 comments

  • mmastrac an hour ago

    Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash

    I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experimenting with a two-node DS4 and it's _good_ at some tasks, but it really just spins its wheels when it hits the limit of what it can reason through.

    I can offload mundane/basic tasks to DS4 on two sparks, but I've been pushing it harder on some novel work and it just can't run on its own at all beyond a certain complexity level.

    I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.

    • disiplus 5 minutes ago

      I will give it a try, but from the benchmarks it never exceeds the DS4 flash benchmarks by significant margin and And I feel that the throughput that you will get on those machines or what I'm getting with my local hosted flash will be so much worse that it's not worth it.

    • Aurornis 6 minutes ago

      > I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.

      Agree. It doesn’t even have to be local, using models in this size class through OpenRouter will reveal their limits if you work side by side with Opus level models regularly.

      There are a lot of social media posts about people cancelling their Anthropic or ChatGPT subscriptions after installing a local LLM. I’ve used local LLMs a lot and I spend a lot of time with frontier models and the difference is still huge. As far as I can tell, the social media posts about local LLMs replacing frontier models are either wishful thinking, engagement bait, or people who must be working on much simpler projects with a much higher tolerance for slop than I have.

      • disiplus 2 minutes ago

        To be fair, there is no 3 turns that I don't have to jump in into what Opus 5 is doing. There is either some regression or my prompting skills are so much worse now. Flash is not perfect and honestly some things depend on how big context do you keep. So I'm keeping like a really short context with my flash, but it works okay, even though it has a tendency to overthink, and yeah, I run it always in max effort mode.

    • kilroy123 an hour ago

      > get myself four sparks at a decent price

      Wow, if you don't mind me asking. How and where?

      • mmastrac 42 minutes ago

        I bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely.

        They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.

        • bmurphy1976 20 minutes ago

          ~$4000 USD each on Amazon, $175 for the cable.

        • swiftcoder 22 minutes ago

          > it's the only model in the whole lineup that isn't priced insanely

          $4,000 isn't priced insanely? ye gads

          • swatcoder a few seconds ago

            [delayed]

          • a3w 20 minutes ago

            I thought 4000 in sum. No wait, 4000 per, plus tax. Or EUR pricing to similar accord. Ouch.

            • swiftcoder 16 minutes ago

              Yeah, that little cluster costs about the same as a brand-new Dacia Sandero.

          • esafak 8 minutes ago

            Yes, but it was $200 off!

        • cmrdporcupine 12 minutes ago

          I mean, I have the same machine and the pricing is only what it is because it has that 1TB nVME in it instead of larger. nVME prices are insane and have been for months.

          Reality is on a single spark I'm constantly running out of room and it being an odd size M.2 slot it's a pain to upgrade. I'm setting up a NAS over RDMA via ConnectX though, that's fun.

    • 0xbadcafebee 11 minutes ago

      If you used the bare API pricing, 1M tokens @ 30% input/70% output/50% cached, you'd pay $0.05805. Even with four discounted sparks, how much are you paying for the same tokens/distribution?

  • sunbum an hour ago

    > with all of this traffic served on Chinese AI chips

    RIP Nivida shareholders

    • Bluestein 44 minutes ago

      This is the takeaway here: That's how they have been serving it at scale as Ox-Alpha. This is a definitional moment.-

      Further quote:

      "Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale."

      https://z.ai/blog/glm-5.3-flash

    • dannyw 28 minutes ago

      Another self-inflicted own courtesy of US government policy.

      While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.

      • ignoramous 18 minutes ago

        The export controls were revoked before it triggered Chinese protectionism: https://www.silicon.co.uk/e-innovation/artificial-intelligen... / https://archive.vn/B2pah

        • mlinsey 6 minutes ago

          Revoked or not, just ever having those controls signals to the Chinese ecosystem that you're not necessarily a reliable supplier (Would you trust US export policy to remain stable for the next ~decade given the state of US politic?) and to the Chinese government just how strategically important you see these components.

          This isn't the kind of thing you can hash out in public and go back and forth on. Once you put it out there, the other party will take steps to make sure they don't have to rely on us in the long run.

        • re-thc 6 minutes ago

          > The export controls were revoked before

          Zai is on another "export control" list outside the broader 1. Doesn't help.

    • Aurornis 4 minutes ago

      Ox Alpha is a smaller model and it was running very slowly. Chinese AI accelerators are coming along, but nVidia’s lead is huge.

    • redox99 8 minutes ago

      Not really a brag: it ran like shit. Very slow (~20tps, VERY high latency) and it would timeout all the time.

      I'm sure the chips are fine, but they clearly didn't have enough capacity for the demand they had (that 100T/day claim was asbolute bs)

      • nchmy 6 minutes ago

        seems unlikely that they'll get nearly as much demand now that it isnt free

    • ChoosesBarbecue an hour ago

      God I wish I could’ve shorted NVIDIA right now

      • browningstreet 43 minutes ago

        It's earnings day for them...

        • re-thc 37 minutes ago

          Which 9/10 times hasn't been great anyway (stock reaction).

          • Bluestein 5 minutes ago

            Of course the release was not coincidental - with the earnings days - I am sure.-

    • ThouYS 30 minutes ago

      yay, I called it! :) (in the other thread)

    • rvz 37 minutes ago

      This is no surprise [0] [1].

      >> "They are already there on open weight models and Jensen knows that it is only a matter of time until China catches up with GPUs or other AI accelerators."

      It is also why Nvidia becoming a bank for other AI companies who are unable to find VCs to fund them isn't really a good thing and that is bearish.

      [0] https://news.ycombinator.com/item?id=49397204

      [1] https://news.ycombinator.com/item?id=49431231

  • revolvingthrow an hour ago

    > 320B total parameters and just 18B active parameters

    This is pretty hefty for a "flash" model, even a 256 GB setup is insufficient at q4 - and q4 is already the worst-but-still-acceptable quant in my experience. The benchmarks look great, especially since GLM tends to be more honest than the average Chinese lab, but you’ll need to splurge to run it at home.

    @edit: so many releases that I forgot to math. This fits just fine in q4, realistically the minimal hardware would be 192gb - so blazing fast on double rtx 6000 pro and usable on 256gb unified memory. You could even go with 5bit quant on 256gb.

    … you’ll still need to splurge, though.

    • colingauvin 42 minutes ago

      That's 160GB-ish for Q4...how is 256 insufficient?

    • dannyw 31 minutes ago

      Looks like the M5 Ultra Studio wait times are going to increase again. Already at 10-12 weeks, I wonder how long it'll go?

      • speedgoose 21 minutes ago

        I guess like the M3 Ultra, at some point normal customers won’t be able to buy it.

  • yipinwong 20 minutes ago

    When reading this type of announcements, always have keen eyes on graphs.

    e.g. "Agent Coding Performance by Effort Level" cuts Y-axis from 0~20.

    - This makes it as if GLM-5.3-Flash made a bigger jump than it claimed as the Y-axis does not increase much (stupid trick used in biz reports)

    I did mention that ox was working ok for me, and having an open-weight comparable to close to SOTA makes it very compelling for me to try it out locally (well, only if I got more VRAM)

    • nchmy 12 minutes ago

      they also conspicuously omitted GPT 5.6 Luna from comparison. It scores lower, but is also cheaper. MiMo 2.5 is not a valid comp at this point

      edit: nevermind. it is there in the artifical analysis scatter plot, but is greyed-out.

      MUCH more interesting is that in that chart, their cost is WAY off. The actual chart shows GLM 5.3 Flash at $0.09, but their chart shows $0.045...

  • TaLiTr 42 minutes ago

    > it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

    From a biased source, but would be big if true. I've had great results with GLM 5.2.

    From their subscription page, the smallest plan gives you about 97M tokens weekly for 5.3 but 292M for 5.3 Flash. Not exactly 10x the limit.

    • wolttam 29 minutes ago

      The recent and slightly smaller DSv4 Flash is also GLM 5.2 equivalent (or close enough)

    • re-thc 38 minutes ago

      > From a biased source, but would be big if true. I've had great results with GLM 5.2.

      It's at least close (even if not better) from the Ox Alpha runs. For the price it's definitely great.

  • cootsnuck 17 minutes ago

    If we fast forward say 5 years, I don't see how we don't end up in world where people (and enterprises) are more savvy with how they use LLMs. Meaning, more models, smaller models, weirder models, more specialized models, etc. And all of it running on a variety of hardware (edge devices, personal computers, on-demand cloud compute).

    I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends are truly to become as useful and ubiquitous as everyone thinks they will, then commoditization is the only option.

  • packetlost an hour ago

    For those who didn't read, this is the identity of the mysterious "Ox Alpha" model

  • mariopt 41 minutes ago

    It's only 320B, local frontier AI is getting closer, sooner than expected.

    • oceansky 29 minutes ago

      Can't come soon enough!

  • garo-pro 35 minutes ago

    > Combined with our latest 30T-token multimodal pre-training corpus [...]

    Is the optimal formula still 20x the amount of model params in tokens for training? Could this mean we're getting a GLM with 1.5t params?

  • rahimnathwani an hour ago

    Related: https://news.ycombinator.com/item?id=49446422

    (281 points, 118 comments)

  • epolanski an hour ago

    I'm starting to think that this whole sanctioning China may motivate and prompt them to do more and better in every field.

    It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.

    • ricardobeat 37 minutes ago

      Starting? This was obvious way back in 2019, when the US decided to give China a little push developing their own silicon industry.

    • himata4113 26 minutes ago

      Well the big problem with china is that they do not respect international law when it comes to technology theft. But that argument is very weak when it appears that a lot of what they do is out in the open for anyone to replicate.

    • esperent an hour ago

      This has been clearly stated as what would happen going back several decades at least.

  • Destiner an hour ago

    from the article, pareto frontier for open source models is completely dominated by GLM now.

    • montroser 21 minutes ago

      Well, it will be interesting to see where Qwen3.8-Flash-Next ends up landing, also released today. These are exciting times!

  • iamsyr an hour ago

    Standard API Pricing for GLM-5.3-Flash (per 1M tokens)

    - Input: $0.15 - Output: $0.50 - Cached input: $0.03

    • Xunjin an hour ago

      Is that cheaper than DS4 flash?

      • nateb2022 a few seconds ago
      • javier123454321 41 minutes ago

        All I can say is that even if it is, I was almost glad to go back to using DS4 Flash. Because 0XAlpha was just so friggin slow to complete a task because of the level of circular reasoning that it would go over and over into, sometimes even returning no output. If I just wanted something done I would switch from a free model to a paid one which is crazy.

        • denysvitali 39 minutes ago

          Tbh it was also slow because it was being hammered by everyone making use of the free tokens

          • javier123454321 25 minutes ago

            Possibly influenced by that, but I believe that is a different issue. I meant the way it processed a request. It went into so many more loops of thinking.

      • swiftcoder 19 minutes ago

        It's even cheaper than DS4's off-peak pricing. Seems like DeepSeek have some stiff competition now

  • swingboy 34 minutes ago

    How much is the “discounted” pricing they mention?

  • kayleykiwi 23 minutes ago

    This looks like it goes hard, can't wait to try it

  • Imustaskforhelp 32 minutes ago

    > To overcome the relatively limited compute and memory capacity of individual chips, we built a dedicated inference engine for this architecture on top of SGLang. Notably, this effort was accelerated by our GLM-5.3-powered infrastructure agent, which assisted engineers in developing and optimizing kernels, diagnosing performance bottlenecks, and improving the serving stack — creating a feedback loop in which the model helped optimize the system serving the model itself.

    > (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.

    It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.

  • toppy 31 minutes ago

    By clicking this link you download some PDF in the background

    • krystofee 31 minutes ago

      Its displayed in the html...