Why we write our own C and C++ inference engines

(localai.io)

12 points | by eatonphil 2 days ago

4 comments

  • dennis16384 40 minutes ago
    I had a similar success with Model2Vec static embedder and NER inference (both GGUF, compiled for WASM), ported to plain C from ONNX Runtime.

    Wasm size from 30Mb to 300kb and 1.5x speedup. It's definitely worth it for performance or distribution size.

  • stephbook 1 hour ago
    Should have started with writing your own blog posts.
    • nnevatie 41 minutes ago
      Came here to say the same. Really tiring to read these slop-infested posts, where everything has the “right shape”.
    • altmanaltman 28 minutes ago
      I went through the post because of your comment but it really doesn't look like AI slop. Can you please share why you feel like its slop and not written by a human? I can also say "should have started writing your own comments" to you and its unfalsifiable. Blanket accusations with no proof is not a good move really.
  • adithyassekhar 29 minutes ago
    What you get: X is the A, Y is the B.
  • federicoTXTS 1 day ago
    [flagged]