45 comments

  • Luker88 a minute ago

    Currently I am running llama-cpp with `Qwen3.8-Flash-Next-UD-IQ3_XXS` on an old ryzen 8845HS with 96G of ram (and no dedicated graphics card) at 7tk/s and ~60tk/s filling, max ~120K context window.

    Surprisingly useful as long as you can leave it running a couple of hours at the very least.

    While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.

  • kamranjon 30 minutes ago

    Dwarfstar already supports this, curious how it compares, but I use the q4 quant daily and it works really well.

    https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...

  • deadbunny an hour ago

    > Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.

    And I thought piping to bash was bad

    • gchamonlive an hour ago

      Piping to bash is definitely worse because there is no plan mode in bash. Agents also normally don't execute anything transparently, at worst you'll see it doing something weird in the logs.

    • snehesht an hour ago

      Yeah, I was surprised at first then had to dig through setup.py and setup.sh files to figure out.

  • snehesht 2 hours ago

    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    https://huggingface.co/Qwen/Qwen3.8-Flash-Next

    • roscas 29 minutes ago

      Coder version with 30t/sec on a Ryzen 3600x with 48GB of RAM with a nvidia 3080.

      This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.

      I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.

      I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.

      This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.

      • StumpChunkman 6 minutes ago

        How much VRAM on your 3080? I've got an early 10gb model. I've been thinking of exploring local coding models, but everyone seems to use much better GPUs than I have access to. Yours is one of the first I've seen with maybe similar hardware on some level.

    • thatsabadlook an hour ago

      Why is this surprisingly well? It's 2.5x faster than anthropic models, you have data sovereignty, privacy,and that's a strong model. Sounds like a best case scenario to me

    • proc0 an hour ago

      Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.

      • incognito124 an hour ago

        Qwen 3.8 flash next is way better than 27B. It's so good I dont even use claude anymore

        • mickeyp an hour ago

          I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.

          It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.

        • snehesht an hour ago

          Yeah I agree, I'm running it with Pi didn't notice much difference compared to lower tier models and the speed, of course.

          • nicce an hour ago

            I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.

      • thatsabadlook an hour ago

        Significantly better for both performance and real world use case. 3.8 27b is a good small model. This is a good model.

        • geye1234 39 minutes ago

          I find 27B more accurate -- maybe because I'm running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.

          • PcChip 15 minutes ago

            Spelling mistakes?

            What inference engine are you using for flash next?

  • nialv7 21 minutes ago

    There are so many AI generated inference engine for local models now, each of them are generally narrower but they are all faster than llama.cpp. Maybe llama.cpp needs to rethink their strategies...

  • prettyblocks an hour ago

    I've been playing with this on a 3090 and it FLIES. Does a pretty good job too on the tasks I've thrown at it (php code base security audits).

  • ryan_glass 27 minutes ago

    Anyone know how it compares to GLM 5.3 for real world use?

  • Tepix 37 minutes ago

    Q2 quantization. Not interested.

    • tcdent 13 minutes ago

      All of these projects targeting low spec systems and "100 tok/s" are the same 2 bit quant without much else. Conveniently none of them include any mention of accuracy in their published numbers. 4 bit is the floor.

  • Neywiny 17 minutes ago

    I don't like that some configuration is fine via arguments and others by environment variable. I've noticed LLMs like doing this. And even more, like hallucinating such things. To me the advantage of AI coding is that the boilerplate of command line arguments and passing them around becomes trivial instead of tedious.

  • gdevenyi an hour ago

    I had this working with the FreeToken inference engine a month ago when they launched.

    https://github.com/FlashML-org/FreeToken

  • hypfer an hour ago

    Is these another one of those repos where it turns out that claude decided to quant the KV cache to q4 or smaller?

    The Readme doesn't say, but it's all AI generated, so..

  • esafak an hour ago

    Has anyone calculated the effective intelligence of these quantized models?

    • nsagent 38 minutes ago

      See this recent paper: Quantization Degradation in Large Language Models: A Signal–Noise Perspective [1].

        We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
      
      This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.

      [1]: https://arxiv.org/abs/2608.08188

      • merbanan 31 minutes ago

        I created a pruned experts model of the q2 quant, while it gave good performance on limited hardware there was severe quality degradation.

    • mkl an hour ago

      There's some info in the README, including:

      > Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.

      https://github.com/Niko1221/Strata#which-model-should-i-pick

      • nicce an hour ago

        I wonder how this Coder compares to Qwen 3.8 27B. Can it be really better since they are competitive for same memory requirements?

      • nisarg2 25 minutes ago

        92% is halfway to 99%

        Holds up pretty well

      • javier2 an hour ago

        ok that is getting interesting!

  • panny 44 minutes ago

    I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM.

    • somenameforme 21 minutes ago

      The card in question here had an initial MSRP of $1600. It's been bumped up by the market, probably because it turned out it's nice for things like this, but it's hardly in the 99% can't afford domain, especially if you're using it to replace a never-ending rent at which point it will pay for itself very rapidly, especially for heavy LLM users.

      In any case, we've gone from requiring supercomputers, to requiring very high end computers, to requiring $1600 video cards. It's tracking the exact same path that image rendering systems took (which if you haven't been keeping up there, now run excellently on pretty much any plain old computer), and we'll probably be there within a couple of years if not much sooner.

    • MaxikCZ 11 minutes ago

      The reason models running on low vram are not talked about enough is because they are just not worth it. Qwen3.8 27b changed that, but even 24gb vram is too low for it. Running better model faster at 12gb vram is where its now at, and thats why you see people talkin about it

    • MrDrMcCoy 40 minutes ago

      Ternary Bonsai 2 might be for you.

  • quietFalcon an hour ago

    Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context?

  • 0xbadcafebee 41 minutes ago

    Lol, sure, if you quant it to hell (Q2) it'll go real fast...

    They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.

    It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.

    • sigbottle 12 minutes ago

      It's interesting though that Q4 seems to be enough, is there a reason that 4 bit floats are good enough for inference?

      • amelius 3 minutes ago

        3 is the magic number, and 4 > 3.

        (seriously, nobody knows why any of this works; it's just a matter of trying)

      • MaxikCZ 7 minutes ago

        New models are trained with 8/4bit quantization in mind. Going from "native" 8 to 4 isnt as big of a step as going from 8 to 4 if native is full bf16.

    • snehesht 14 minutes ago

      You're right but for simple use cases its useful. Someone pointed to ds4 + qwen3.8 with Q4_K will try that out.