36 points | by graham33 2 hours ago ago
8 comments
Huge plus to anyone interested in the space to check out what Graham built here!
If anyone is also interested on Nix/CUDA/Capital Markets/Flox, we recently did another case study in the space - https://flox.dev/blog/deploying-hardened-flox-nvidia-cuda-st...
This is incredible. I have a Jetson lying around and will try to it out on this. I use it to play with vision models, not LLMs, and have been wanting a better way to manage the machine.
Thanks for sharing, saving this for when I get a DGX Spark
Been running this on a few Asus GX10 machines with k3s on top, it’s been great. I’m running the new deepseek.
Thank you for your work!
What quant are you using, and what tps are you getting with K3?
The FP8 version from DeepSeek themselves [0], around 1800 tps prefill and 45 tokens per second decode.
I’ve been running a custom VLLM image with b12x as well as nvfp4_ds_mla.
I would say it’s quite fantastic in day to day, I use it mostly in Hermes and sometimes for coding.
I have qwen 3.6 27b in an rtx 6000 pro as well so I use that as a workhorse in pi with DS as a reviewer/planner.
[0] https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
Edit: I think you may have misread my post. k3s is NOT kimi k3, and I did mention I was running deepseek.
This has been amazingly helpful for managing my DGX Spark! Thank you for all your time and effort into this project!
Good to hear, thanks!
Huge plus to anyone interested in the space to check out what Graham built here!
If anyone is also interested on Nix/CUDA/Capital Markets/Flox, we recently did another case study in the space - https://flox.dev/blog/deploying-hardened-flox-nvidia-cuda-st...
This is incredible. I have a Jetson lying around and will try to it out on this. I use it to play with vision models, not LLMs, and have been wanting a better way to manage the machine.
Thanks for sharing, saving this for when I get a DGX Spark
Been running this on a few Asus GX10 machines with k3s on top, it’s been great. I’m running the new deepseek.
Thank you for your work!
What quant are you using, and what tps are you getting with K3?
The FP8 version from DeepSeek themselves [0], around 1800 tps prefill and 45 tokens per second decode.
I’ve been running a custom VLLM image with b12x as well as nvfp4_ds_mla.
I would say it’s quite fantastic in day to day, I use it mostly in Hermes and sometimes for coding.
I have qwen 3.6 27b in an rtx 6000 pro as well so I use that as a workhorse in pi with DS as a reviewer/planner.
[0] https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
Edit: I think you may have misread my post. k3s is NOT kimi k3, and I did mention I was running deepseek.
This has been amazingly helpful for managing my DGX Spark! Thank you for all your time and effort into this project!
Good to hear, thanks!