Running Kimi K3 on a M1 Max

(github.com)

85 points | by tito 3 hours ago ago

64 comments

  • antirez 3 hours ago

    SSD streaming on an M5 Max 128GB: https://x.com/antirez/status/2082136334160818528

    Soon decent speed across two Mac Studios with 512GB of RAM.

    • whatsThisBtn4 2 hours ago

      0.3tx per second is decent speed?

      And it gets worse with every token.

      • smallerize 5 minutes ago

        That's not what he said. He said with 2x the hardware it will be faster.

    • xfour 2 hours ago

      Cool stuff. Do you have a the hardware and a way to bridge the compute? Or just hopeful?

    • tito 3 hours ago

      wow cool! I like watching the new models come out and how they end up crammed in to run on local machines. I learned what mxfp4 is thanks to this latest Kimi release - although it sounds like it means that there's less room for compression in the model compared to others.

    • api 34 minutes ago

      I’ve wondered for a while: given the lower cost of SSD per GB could you build a very wide RAID0 style striped array of SSDs (maybe one per slot) to get almost RAM like read speeds?

      To really go fast you’d probably have to do PCB layout and do like 256 or 1024 chips in parallel with a fast SRAM aggregation buffer feeding a GPU or TPU rig.

      Or could you do the same with custom layout of cheap slower RAM?

      I wonder if anyone is doing this? You would flash in a model and then just run it. It would need RAM for context but much less of it.

  • Azantys 3 hours ago

    0.01 tk/s is unusable for anything, you would wait a whole day for just 1000 token of output, what is the point of projects like this?

    • SXX 3 hours ago

      It is fun.

      Also its answering the question of what gonna happen if you wake up tomorrow and datacenters are gone. Or internets are gone.

      Some people on our globe live in countries with no internet whatsoever. Of course most of them dont have Macbook with 64GB RAM either, but it's much much easier to get than internet connection or rack of GB200.

      SOTA LLMs are efficiently compression of all the knowkedge humanity has built. Having ability to run it at home to extract said knowledge is important no matter the speed.

      • whatsThisBtn4 an hour ago

        Even if somehow all the data centers are gone, it's still uselessly slow.

        But also that's a pretty extreme hypothetical. Imagine the polymarket on that.

    • ggm 2 hours ago

      So you subscribe to the belief we won't in future find mentalism in other galaxies or solar systems which operate on mechanisms we don't understand and think v e r y s l o w w w w w w w l y ?

      (note. I am not a believer in AGI)

      "useful" is highly contextual. The clock of the long "now" is not useful in the sense you mean, to synchronise your wristwatch. I'm still glad it exists.

      • tito 2 hours ago

        Are there any well thought through stories about what this would look like? For example, I'm thinking about like nutrient flow, decision making, energy input, gravitational force, things like that seem to govern the value and speed of intelligence.

        • indiv0 42 minutes ago

          Hard to avoid spoilers here but Vernor Vinge hits on pretty much exactly this in A Fire Upon the Deep. Though he's interested less in the hard sci-fi aspects of how/why and more on the consequences of it (story-wise).

        • ggm an hour ago

          I think this is a slow version of the quandry behind Quantum Computing: how do you distinguish events from the noise floor? It happens in the quantum context and it would happen in the millenial timeframe completing "operations" which have to be compared to e.g. the stability of orbit around a sun.

        • kirubakaran an hour ago

          For the opposite, "Dragon's Egg" by Robert L. Forward is a fantastic read

    • tito 3 hours ago

      I like seeing the latest and greatest model crammed into new systems to see how it fares. To deal with the speed, one person on reddit suggested using it in an email interface rather than a chat interface.

      • 0xc133 11 minutes ago

        I had my clanker implement this idea in a standalone Rust server that speaks IMAP and SMTP and proxies your emails to an OpenAI endpoint you configure: https://tangled.org/clee.sh/posthorn

        Works in mutt; other MUAs may vary.

      • magicalhippo 2 hours ago

        > one person on reddit suggested using it in an email interface rather than a chat interface

        Kimi Pen Pal. Bring back lettets and postcards. Do OCR, and use one of those 3D printer-like pen plotters write the model output as a letter.

        Challenge would be automating the opening and OCR preparation, and the folding and mailing of the return letter. But given it's done commercially it should be possible.

      • embedding-shape 2 hours ago

        Email would indeed be fitting for K3 running on a M1 Mac, as it'd take days/weeks to receive a response, which matches with my real-world emailing experience pretty well.

      • hugopuybareau 3 hours ago

        Love the email idea

        • tito 3 hours ago

          Having the right type of interface makes a huge difference.

          It reminds me of when Willow Garage chose to name their bot the TurtleBot, because if they named it anything else, people would think it was fast and capable. But when they called it Turtle Bot, people just kind of liked it and were satisfied with what it did.

          At the level of Kimi 3, I probably can code only about 1,000 good tokens per day, too. (thankfully coding isn't my job)

    • anigbrowl 41 minutes ago

      'Large Language models? They can barely produce gibberish sentences, what would this tech ever be useful for?'

      - bunch of people only ~4 years ago

      • Azantys 29 minutes ago

        I questioned the speed not the output quality, that is another discussion

        • bibstha 10 minutes ago

          It applies to speed too. The project paves way for more optimization at many layers overtime.

    • whatsThisBtn4 an hour ago

      I can only guess some post purchase remorse.

      Need to justify buying an expensive rig that doesn't do what you expected.

      Specifically thinking the people they could do something AI with cpu, and realizing it isn't feasible. Happened at my fortune 20 company. They had to get approvals and ofc it was useless. Plenty people tried to explain, but they were the principle engineer, and out ranked everyone.

      "It's not going to work", the topic changed, and we never spoke about it again.

    • winstonp 2 hours ago

      16 tokens / s is not nothing.

      • throwaway219450 2 hours ago

        It's the other way around due to poor framing, it would be much easier to compare if you [the repo] said 0.02 tps.

      • whatsThisBtn4 an hour ago

        Confusing numbers everywhere.

        16tk/s... Then 3 tks per minute. Then someone else posted 0.3tk/s.

      • Azantys 2 hours ago

        ? Readme says 60-70s per token

  • ALLTaken 2 hours ago

    Exactly my machine 64GB M1 Max So happy about this! ♡

    idk how people access (soldout) and even afford 512GB RAM MacStudio's. Isn't it $40k or so?

    • hmokiguess 2 hours ago

      You're lucky, mine is the M1 Pro 16GB

    • teaearlgraycold 32 minutes ago

      I just checked eBay. There's an insane price difference between used and new. $5k vs $40k.

  • mips_avatar 2 hours ago

    Would be interesting to see how fast it would be on 4x mac studio 512gb machines.

  • mindwok 34 minutes ago

    Does anyone else feel like the writing is on the wall for a future of local models? Spamming data centres everywhere, powering them, having to commit insane capital to hardware, all the effort to serve inference over a network reliably - when here we are with a frontier model nearly running on a laptop.

    Local AI on your device seems like a much more likely future to me than datacenters in space. For inference at least, training is another story.

    • nomel 17 minutes ago

      0.01 tk/s on an M1 Max is not "nearly". This is completely unusable, and in no way cost effective.

  • nlessard 3 hours ago

    Anyone who knows the state of NVMe hardware more than me know if this would obliterate the lifespan of your drive? Seems like the biggest limitation to me (some people are probably fine with letting their Macs churn over the weekend).

    • trollbridge 2 hours ago

      No problem at all to read data over and over. In fact, LLM weights are a great candidate for low-quality flash that can't handle a lot of write cycles, and you want a large amount of storage cheaply...

    • addaon 3 hours ago

      Reads are not generally life-limiting for flash. (Well, no more so than power-on time in general. You still have aging mechanisms like electromigration, but these are orders of magnitude slower than write-induced damage.)

  • acmnrs 2 hours ago

    The title should probably be edited to specify "M1 Max" instead of "M1 Mac". You aren't running K3 on a base M1 anytime soon. Either way, still a very impressive project.

    • tito 2 hours ago

      Done, Mac -> Max

  • jrhizor 44 minutes ago

    Super cool, and I appreciate the upfront speed disclaimer

  • hmokiguess 2 hours ago

    Now set it up with an agent and a permanent `/goal` to say it cannot stop until it has solved for speed, then leave it on and livestream so we can all see when it becomes exponential. Could have the Eternal Jukebox playing in the background!

    • Anoian 2 hours ago

      That sounds like the most boring exciting livestream of all time.

  • denysvitali 2 hours ago

    60s/token - if only there was a way to drop that "s" this would be amazing

  • walrus01 40 minutes ago

    tokens/second, no, more like seconds/token

    • ncallaway 21 minutes ago

      The link beat you to it:

      > It is not fast — about 16 seconds per token on our M1 Max

      • walrus01 14 minutes ago

        At least it's not minutes/token?

  • tjwebbnorfolk 2 hours ago

    > ~60–76 s/token

    I don't know if I'd call this "running"

    • sermah 2 hours ago

      Had the same feeling when I first saw min/km units in some (human) running context.

      UPD: I know it's not the same at all, just the reversal of units that gets me

    • tito 2 hours ago

      I commented similarly below, but as a terrible programmer, I probably perform about 1 minute per token too (at Kimi 3 level). It puts into context how I think about intelligence

      • brokencode 2 hours ago

        Maybe in terms of code produced, but one token is only a fragment of a thought for an LLM.

        It’d be like thinking as slowly as Ents talk to each other in Lord of the Rings.

        • tito 2 hours ago

          Oh, is that how it works? So, when somebody says a model is running at X tokens per second, it means that the thinking process is running at that, and output tokens are much lower then? Thanks to the explanation.

          • brokencode an hour ago

            It’s all just tokens to the model. Whether it’s thinking tokens or output tokens, they take the same amount of computation to produce. The only difference is whether the token is displayed to the user.

  • als0 3 hours ago

    Says it requires a 2TB disk? Must it be internal NVMe?

    • Fergusonb 2 hours ago

      You can use an external drive if it's mounted as a writable volume. I would make sure it's fast, maybe thunderbolt 3/4/5 enclosure with a fast drive.

      • Xeoncross 2 hours ago

        FYI, Thunderbolt 3 NVMe enclosures will be 10Gbps and 4 will be 40Gbps which is a big difference (and you'll notice it in the pricing as well)

    • tito 3 hours ago

      The Github specifically mentions an option to stream it from an online host. It's extremely slow.

  • lostmsu 3 hours ago

    under 0.02 tok/s

  • piterrro 3 hours ago

    Will it fit on ESP32??

    • embedding-shape 2 hours ago

      Better questions, how many ESP32s would it take to reach 1 tok/s decoding speed with K3?

    • tito 3 hours ago

      1 button Kimi morse code interface

  • onesandofgrain 2 hours ago

    Cool gimmick

  • brcmthrowaway 2 hours ago

    Why not train another smaller LLM to give the same answers as Kimi K3?

    • dkarras 6 minutes ago

      why not zip the entire internet to 1MB so everyone can have a copy? because it is not possible - we don't know if it is possible. I mean we know it is impossible, but we don't know if it is possible to do it with acceptable quality loss.