5 comments

  • dlcarrier an hour ago

    From what I've seen, Vulkan adds a lot of overhead on Intel hardware.

  • PcChip 3 hours ago

    I didn't see any benchmarks against vllm, sglang, exllama, etc

    • rancor 3 hours ago

      Since this is basically a wrapper around libllama.so, I would assume that the performance is roughly the same as llama.cpp upstream.

  • peddling-brink 2 hours ago

    > llama.cpp via Vulkan (AMD / Intel / NVIDIA) or CPU fallback

    I got excited about someone paying attention to intel. Oh well.

    • kamranjon a few seconds ago

      llama.cpp sycl and vllm xmx work is pretty incredible right now - you just gotta build it with some extra flags