4 comments

  • PcChip 6 hours ago
    I didn't see any benchmarks against vllm, sglang, exllama, etc
    • rancor 6 hours ago
      Since this is basically a wrapper around libllama.so, I would assume that the performance is roughly the same as llama.cpp upstream.
  • dlcarrier 5 hours ago
    From what I've seen, Vulkan adds a lot of overhead on Intel hardware.
  • peddling-brink 5 hours ago
    > llama.cpp via Vulkan (AMD / Intel / NVIDIA) or CPU fallback

    I got excited about someone paying attention to intel. Oh well.

    • wronglebowski 2 hours ago
      What hardware do you have? I’ve been playing with a 258V and OpenVINO has come a longggggg way.
      • peddling-brink 1 hour ago
        Two arc b60s. The intel vllm build is getting me ~15t/s decode with heavy context using qwen3.8 27b.
    • kamranjon 3 hours ago
      llama.cpp sycl and vllm xmx work is pretty incredible right now - you just gotta build it with some extra flags
      • peddling-brink 1 hour ago
        Llama would be nice for the ggufs. Any specific flags or tutorials I should look at?
  • shayanjavadi 4 hours ago
    [dead]