Show HN: Shoehorn – Quantize any model down to run on your machine

(notactuallytreyanastasio.github.io)

34 points | by rhgraysonii 4 hours ago ago

3 comments

  • hmokiguess 3 hours ago
  • jaylane 23 minutes ago

    tried it out but based on the model sizing result i got i got an insufficient memory error when the server started running

  • mbuchel-hn 4 hours ago

    does this work similar to airllm? i am wondering how it would handle something like quantizing kimi k3 on a budget of 8 gbs, or is that something you are not attempting to solve yet?