it IS cheap (once the weights are released) and it will only become CHEAPER. For around $3700 a month (via loan purchased hardware + energy cost) you can run around 32 concurrent instances of kimi k3 which can generate nearly a 6.9 billion tokens a day.
This is napkin math since I'm mostly just extrapolating from glm 5.2 by assuming it's twice as heavy to serve in every single measurement, but I believe you can easily achieve 2500tok/s aggregate compared to 4500tok/s and up to 8000tok/s for glm5.2.
with nvidia r100 you are likely going to be able to push that number even higher while the cost of hardware appears to be relatively the same, so far I am seeing 21% premium from supermicro which is twice as fast and has nearly twice the vram.
I was actually shocked to see that much improvement on GLM 5.2. I am getting pretty great rates in GLM5.2 right now and I'm extremely happy with the output. I have found Kimi to be generally a little slower but noticeably better at architecture and structuring things. I would have thought is for sure currently more expensive than GLM5.2, particularly with subscriptions etc, but I'm excited for this to decrease further.
Agreed, Kimi is cheaper for coding - I say that explicitly in the post too. However I'd have to disagree with you on the "office task" front.
General office work is one of the big frontiers the labs are pushing on, and it's part of how they're justifying the value proposition to enterprise customers. It's also accounts for a big portion of the spend on RL; tasks/environments designed to train agents to navigate Slack or Salesforce. If you're Anthropic pitching Claude to a bank (taking an example I'm familiar with), coding probably accounts for ~20% tops of the workforce, and it doesn't drive direct revenues. The 'agentic coding bump', but for all your analysts, traders, and wealth managers, would be a much more attractive prospect.
I don't disagree that coding is the most successful use case so far (and probably more relevant to a HN audience). But I think the future of the labs is also contingent on them making progress on more general white collar work. I suspect that's why the Opus 5 release blog lists 3 coding benchmarks (FrontierBench, DeepSWE and FrontierCode) to 3 or 4 more general ones applicable to office work - depending on how you slice it (GDPVal, AutomationBench, Legal Agent Benchmark, BrowseComp).
Okay, I think that's fair, but I'm not convinced there's anybody actually doing large amounts of compute on office tasks? Do you know anybody? Can you point to anybody publicly documenting this? Can you can you point to any specific workflows where fable is being used in lieu of more basic models?
Even if you provide exceptional answers for all of these I still think it is disingenuous at best to ignore coding tasks in writing this. I have to assume coding is 90% of the use cases for the frontier.
You've made a case that the labs need these customers. You haven't made a case that the labs have these customers.
No I think you're right that the amount of compute spent on office work is lower than coding - although I don't have any sense for the right share. The best source I could find was an OpenAI report [1] which mentions that ~66% of enterprise token generation is via Codex, which I would expect to skew entirely towards coding. But it's hard to say how the remainder is split, what proportion is 'frontier', or whether it's representative for Anthropic.
On your questions - I've spoken to a number of execs and seniors behind closed doors but nothing public I can point to. Anecdotally, I've spoken to senior leaders at banks spending billions of tokens on one-off tasks like prepping execs for earnings calls or piloting end-to-end agent workflows for specific use cases (but mostly piecemeal/one-off).
On Fable, financial analysts I know are using it to produce research docs, models and decks - I hear that it's a big improvement for these tasks. This lot have been blindsided by the spend growth [2], the same as for coders in enterprise (e.g. Uber blowing annual budget in 4 months [3]), so I do think there's appetite and budget for a capable, cheaper open model - but Kimi does not obviously fill that role the way it might for coding. That said, I still largely agree with you - where coding has seen a broad deployment across software development, most of the office work stuff is still fairly piecemeal and certainly lower compute-spend.
I think it's fair to say I could've focussed on coding more rather than taking AA's benchmark distribution as representative - perhaps a more balanced title would be "Kimi K3 is not cheap across the board"? I guess there's also some ambiguity about what 'cheap' means - as I said elsewhere in this thread, I think when some people talk about the price of Chinese models, they imagine Deepseek competing with o1 for 1/20th of the price. Even though it is better priced for coding, Kimi isn't Deepseek-level cheap.
I do, however, think you could debate whether coding will remain at >50% total token usage going forwards - big enterprises are hunting for ways to get value out of LLMs, and the labs are investing a correspondingly large amount in generating demonstrations and RL environments to get the models up to par.
It's also possible that Chinese labs will shift focus to white collar applications now they've demonstrated a commanding lead on cost efficiency for coding, so, I mean who knows - it'll be interesting to get some detail when Anthropic IPOs.
Sorry for the long reply! Appreciate it's quite meandering...
K2.6 is cheaper than GLM5.2 (at least on DeepInfra) and I've found it works as good as Sonnet for my purposes. Both tend to think themselves into circles a bit and aren't super token efficient, but I've found GLM5.2 much worse on this count making K2.6 even cheaper than the per token price would make seem.
This article seems premature to post. Right now, the price is arbitrarily set by a single provider. Why wouldn't Moonshot collect extra revenue during this exclusivity period when they knew there would be hype?
The model weights are supposed to release tomorrow.
Over the next several weeks, I would expect competition among open weight providers to drive down the cost, as I've seen happen with other open weight model releases.
I don't mean to imply that Kimi is not at all cheaper than U.S frontier models. I more wrote this because I believe - since Chinese LLMs entered the public consciousness via DeepSeek R1, which was genuinely ~20x cheaper than o1 - there's a bit of a halo effect around Chinese models which causes people to overestimate the scale of the discount. And relative to that price anchor, Kimi is less extraordinarily cheap.
At the moment Kimi is ~10% cheaper than GPT-5.6 on the AA benchmark, and as you say that could go down to 20-30% cheaper (although I don't know how inference provider discounts play out on real world usage once you account for quantisation etc...). I'm not trying to suggest that that's nothing, but I do think some of the people driving the Chinese AI discourse would have a harder time pitching their conclusions if they were saying "this new Chinese model is 10% cheaper on some tasks, and it might get another 20% cheaper in the future".
But if you compare to Anthropic's models? The cost difference is huge. Anthropic is clearly concerned that people are realizing they are expensive, since the Opus 5 blog post dedicated a lot of time to talking about how cheap the model was compared to the competition... but this doesn't hold water when I haven't seen any independent benchmarks claiming Opus 5 is cheaper than GPT-5.6-Sol, even if it is supposedly closer.
GPT-5.6-Sol is pretty competitively priced, but not all American frontier models are, and even 10% to 30% is still significant for any commodity that's as fungible as frontier models often are.
> as you say that could go down to 20-30% cheaper
I never said anything about 20% to 30%. We don't know how much it actually costs to host this model yet, and that will determine the final price. It could be just a little less, or it could be a lot less.
> once you account for quantisation
There will be no need to account for quantization. Kimi models have been 4-bit only since at least K2.5. They don't release or serve models in higher precision than that. This isn't one of those situations where LLM inference providers are debating between serving 16-bit, 8-bit, or 4-bit, and I have never seen a publicly hosted, paid model that was hosted in less than 4-bit, even if hobbyists will use sub-4-bit quantizations sometimes locally.
It's cheap because it won't refuse random tasks. You can't get rid of nannying at any price beyond training your own model, and relative to that, K3 is cheap.
A lot hinges on what happens tomorrow. I'll believe they'll open the weights when I see the files appear on HF (and when somebody with 24 RTX6000s or whatever reports that they are indeed as good as the closed version.)
In my testing, I'm finding it more expensive than Opus 4.8/5 and GPT 5.6 Sol at API rates, because it chews so much. And, their plan (at least the $19 tier) is much less generous than the ChatGPT $20 plan, like an order of magnitude less, it's basically a demo not a useful amount of usage.
Yeah can't recommend their $19 plan, only took me a day and a half to hit the weekly usage cap. The $39 plan has 5x the limits so I recommend getting that instead. Having now tested it for a few days, it is the first of the chinese models that actually feels comparable to Opus-tier models
it IS cheap (once the weights are released) and it will only become CHEAPER. For around $3700 a month (via loan purchased hardware + energy cost) you can run around 32 concurrent instances of kimi k3 which can generate nearly a 6.9 billion tokens a day.
This is napkin math since I'm mostly just extrapolating from glm 5.2 by assuming it's twice as heavy to serve in every single measurement, but I believe you can easily achieve 2500tok/s aggregate compared to 4500tok/s and up to 8000tok/s for glm5.2.
with nvidia r100 you are likely going to be able to push that number even higher while the cost of hardware appears to be relatively the same, so far I am seeing 21% premium from supermicro which is twice as fast and has nearly twice the vram.
It feels like this "Kimi is a token hog" meme is 100% astroturfed by Anthropic. It's cheap. Believe your own eyes.
I mean at this point their very existence depends on it so I’m not sure if I’d be surprised
Not to mention -
If you switch the view to "coding tasks" on this website:
So it's pretty dang cheap lol. Nobody is using frontier inference for "office tasks".Right? The availability of this being in the article that’s pushing the opposite narrative is like…what?
I was actually shocked to see that much improvement on GLM 5.2. I am getting pretty great rates in GLM5.2 right now and I'm extremely happy with the output. I have found Kimi to be generally a little slower but noticeably better at architecture and structuring things. I would have thought is for sure currently more expensive than GLM5.2, particularly with subscriptions etc, but I'm excited for this to decrease further.
Agreed, Kimi is cheaper for coding - I say that explicitly in the post too. However I'd have to disagree with you on the "office task" front.
General office work is one of the big frontiers the labs are pushing on, and it's part of how they're justifying the value proposition to enterprise customers. It's also accounts for a big portion of the spend on RL; tasks/environments designed to train agents to navigate Slack or Salesforce. If you're Anthropic pitching Claude to a bank (taking an example I'm familiar with), coding probably accounts for ~20% tops of the workforce, and it doesn't drive direct revenues. The 'agentic coding bump', but for all your analysts, traders, and wealth managers, would be a much more attractive prospect.
I don't disagree that coding is the most successful use case so far (and probably more relevant to a HN audience). But I think the future of the labs is also contingent on them making progress on more general white collar work. I suspect that's why the Opus 5 release blog lists 3 coding benchmarks (FrontierBench, DeepSWE and FrontierCode) to 3 or 4 more general ones applicable to office work - depending on how you slice it (GDPVal, AutomationBench, Legal Agent Benchmark, BrowseComp).
Okay, I think that's fair, but I'm not convinced there's anybody actually doing large amounts of compute on office tasks? Do you know anybody? Can you point to anybody publicly documenting this? Can you can you point to any specific workflows where fable is being used in lieu of more basic models?
Even if you provide exceptional answers for all of these I still think it is disingenuous at best to ignore coding tasks in writing this. I have to assume coding is 90% of the use cases for the frontier.
You've made a case that the labs need these customers. You haven't made a case that the labs have these customers.
No I think you're right that the amount of compute spent on office work is lower than coding - although I don't have any sense for the right share. The best source I could find was an OpenAI report [1] which mentions that ~66% of enterprise token generation is via Codex, which I would expect to skew entirely towards coding. But it's hard to say how the remainder is split, what proportion is 'frontier', or whether it's representative for Anthropic.
On your questions - I've spoken to a number of execs and seniors behind closed doors but nothing public I can point to. Anecdotally, I've spoken to senior leaders at banks spending billions of tokens on one-off tasks like prepping execs for earnings calls or piloting end-to-end agent workflows for specific use cases (but mostly piecemeal/one-off).
On Fable, financial analysts I know are using it to produce research docs, models and decks - I hear that it's a big improvement for these tasks. This lot have been blindsided by the spend growth [2], the same as for coders in enterprise (e.g. Uber blowing annual budget in 4 months [3]), so I do think there's appetite and budget for a capable, cheaper open model - but Kimi does not obviously fill that role the way it might for coding. That said, I still largely agree with you - where coding has seen a broad deployment across software development, most of the office work stuff is still fairly piecemeal and certainly lower compute-spend.
I think it's fair to say I could've focussed on coding more rather than taking AA's benchmark distribution as representative - perhaps a more balanced title would be "Kimi K3 is not cheap across the board"? I guess there's also some ambiguity about what 'cheap' means - as I said elsewhere in this thread, I think when some people talk about the price of Chinese models, they imagine Deepseek competing with o1 for 1/20th of the price. Even though it is better priced for coding, Kimi isn't Deepseek-level cheap.
I do, however, think you could debate whether coding will remain at >50% total token usage going forwards - big enterprises are hunting for ways to get value out of LLMs, and the labs are investing a correspondingly large amount in generating demonstrations and RL environments to get the models up to par. It's also possible that Chinese labs will shift focus to white collar applications now they've demonstrated a commanding lead on cost efficiency for coding, so, I mean who knows - it'll be interesting to get some detail when Anthropic IPOs.
Sorry for the long reply! Appreciate it's quite meandering...
[1] https://cdn.openai.com/pdf/5d1e1489-21c0-43e4-9d42-f87efdbf0...
[2] https://www.reuters.com/business/finance/australias-cba-flag...
[3] https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-c...
K2.6 is cheaper than GLM5.2 (at least on DeepInfra) and I've found it works as good as Sonnet for my purposes. Both tend to think themselves into circles a bit and aren't super token efficient, but I've found GLM5.2 much worse on this count making K2.6 even cheaper than the per token price would make seem.
I find them to be about comparable, but I use them both for coding tasks and I'm happy with each. I like K3
This article seems premature to post. Right now, the price is arbitrarily set by a single provider. Why wouldn't Moonshot collect extra revenue during this exclusivity period when they knew there would be hype?
The model weights are supposed to release tomorrow.
Over the next several weeks, I would expect competition among open weight providers to drive down the cost, as I've seen happen with other open weight model releases.
That's a very fair critique.
I don't mean to imply that Kimi is not at all cheaper than U.S frontier models. I more wrote this because I believe - since Chinese LLMs entered the public consciousness via DeepSeek R1, which was genuinely ~20x cheaper than o1 - there's a bit of a halo effect around Chinese models which causes people to overestimate the scale of the discount. And relative to that price anchor, Kimi is less extraordinarily cheap.
At the moment Kimi is ~10% cheaper than GPT-5.6 on the AA benchmark, and as you say that could go down to 20-30% cheaper (although I don't know how inference provider discounts play out on real world usage once you account for quantisation etc...). I'm not trying to suggest that that's nothing, but I do think some of the people driving the Chinese AI discourse would have a harder time pitching their conclusions if they were saying "this new Chinese model is 10% cheaper on some tasks, and it might get another 20% cheaper in the future".
Sir, it's 3X cheaper on coding. 3X cheaper in any industry is earth shattering. 10% is significant. 3X is really big.
But if you compare to Anthropic's models? The cost difference is huge. Anthropic is clearly concerned that people are realizing they are expensive, since the Opus 5 blog post dedicated a lot of time to talking about how cheap the model was compared to the competition... but this doesn't hold water when I haven't seen any independent benchmarks claiming Opus 5 is cheaper than GPT-5.6-Sol, even if it is supposedly closer.
GPT-5.6-Sol is pretty competitively priced, but not all American frontier models are, and even 10% to 30% is still significant for any commodity that's as fungible as frontier models often are.
> as you say that could go down to 20-30% cheaper
I never said anything about 20% to 30%. We don't know how much it actually costs to host this model yet, and that will determine the final price. It could be just a little less, or it could be a lot less.
> once you account for quantisation
There will be no need to account for quantization. Kimi models have been 4-bit only since at least K2.5. They don't release or serve models in higher precision than that. This isn't one of those situations where LLM inference providers are debating between serving 16-bit, 8-bit, or 4-bit, and I have never seen a publicly hosted, paid model that was hosted in less than 4-bit, even if hobbyists will use sub-4-bit quantizations sometimes locally.
It's cheap because it won't refuse random tasks. You can't get rid of nannying at any price beyond training your own model, and relative to that, K3 is cheap.
It’s cheap.
I mean you have a chart showing that it’s cheaper than the other models and it also does what I want without argument.
Additionally, I fully expect the frontier labs to continue increasing prices to meet the profit margins they need to to continue existing.
A lot hinges on what happens tomorrow. I'll believe they'll open the weights when I see the files appear on HF (and when somebody with 24 RTX6000s or whatever reports that they are indeed as good as the closed version.)
[flagged]
[flagged]
[dead]
In my testing, I'm finding it more expensive than Opus 4.8/5 and GPT 5.6 Sol at API rates, because it chews so much. And, their plan (at least the $19 tier) is much less generous than the ChatGPT $20 plan, like an order of magnitude less, it's basically a demo not a useful amount of usage.
Yeah can't recommend their $19 plan, only took me a day and a half to hit the weekly usage cap. The $39 plan has 5x the limits so I recommend getting that instead. Having now tested it for a few days, it is the first of the chinese models that actually feels comparable to Opus-tier models