Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

(modelscope.cn)

59 points | by garo-pro 2 hours ago ago

16 comments

  • ddtaylor 14 minutes ago

    I enjoy the Qwen models a lot, but building things on top of them with OpenRouter has been painful.

    OpenRouter does a lot of great work and I really enjoy being able to use different models so easily. I like when a provider is phasing out an older model that still works for my needs and the price is much lower. It seems like such a good win-win.

    However, the problem is that many Qwen models have almost no capacity or is so flaky you literally have to just litter your code with a blacklist/whitelist of providers. OpenRouter has some attempts to solve this, but they don't work. In fact, OpenRouter has a lot of really cool stuff that is documented, but if you read the code it's not yet implemented or isn't actually there yet, which is a shame.

    I tried to get in contact with them at OpenRouter about this and I was interested in working with them in the past, but it's difficult to get in touch with the right people and they are growing very fast. I expect being acquired by Stripe will accelerate those problems in some ways. I have no doubt they will resolve all of these issues eventually and scaling that much that quickly is really hard, so kudos to them, but the road has been pretty lame and taken some wind out of my sails.

    • irthomasthomas a minute ago

      Openrouter was pretty great before prompt caching became common. Now it is extremely expensive for most individual workflows, unless you spend a lot of work customizing router preferences, and then you still get a worse cache hit rate than using the provider directly. I only keep $5-$10 in OR for occasional testing.

  • fcanesin 32 minutes ago
  • pwython 19 minutes ago

    I was already rolling around the idea of a 128GB M5 Max MBP. Now this!

    A 4-bit MLX quant with 128k window should fit perfectly, in the 50-70 tok/s range.

    • sscaryterry 16 minutes ago

      I have a 128GB M5 Max, and it sucks at this stage. 50-70 tok/s might be something...

  • big-chungus4 20 minutes ago

    > We are releasing these architectural improvements ahead of time so that the community can prepare for the upcoming full family of Qwen4 models.

    That gives me hope that "full family" means it will include smaller models like 4B.

  • BrucecarlL 13 minutes ago

    Waiting for the performance report! Ai hope it can beat DS

  • big-chungus4 24 minutes ago

    I hope there is going to be a free endpoint... Unlike 35B-A3B, I am nowhere close to running it locally

  • honestlyranked 31 minutes ago

    Alibaba is giving sleepless nights to the tech giants

  • cogman10 28 minutes ago

    Wow. I wasn't expecting this. I thought they were going to do a 35B model instead.

  • bellowsgulch 6 minutes ago

    [delayed]

  • tw1984 18 minutes ago

    Qwen4 sounds exciting

  • tarruda 2 hours ago

    Can you share the source for the parameter count (125B A6B)? I didn't see it anywhere in the page.

  • mrdoe 29 minutes ago

    lol blocked with dns4eu

    what a joke this resolver has become

  • blurbleblurble 23 minutes ago

    gg