An SLM trained on $8 ESP32-S3

(github.com)

43 points | by pavelai 8 hours ago

8 comments

  • chicken-stew 5 hours ago
    Would probably be better to demonstrate by exsmple how this approach is used to train on sensor data and then use it (as is hinted by the author) instead of acknowledging that the klingon poc is useless.
    • rbanffy 20 minutes ago
      But would it be as cool as a microcontroller that spits out Klingon?
  • runtime_lens 1 hour ago
    Sometimes the proof of concept isn't the product. It's the constraints it exposes that end up influencing more practical systems.
    • wuschel 1 hour ago
      I understand that you are alluring that maximum training-state memory, not parameter count, is the key variable here. So you better start with the smallest model, and go to for highest training data quality, with the outlook of coupling systems together?
  • dannyw 6 hours ago
    Very cool project! Sounds like it was a fun challenge :)

    I wonder when we'll start seeing clusters of ESP32-S3s... not sure how interconnects would go though, but I guess the interconnect wouldn't be the bottleneck anyway.

    • ReactiveJelly 6 hours ago
      Imagine a Beowulf cluster of those
    • fsniper 3 hours ago
      Afaik there are already a few videos on YouTube.
  • prplxd_nihilist 1 hour ago
    > Which is no small thing.

    I think it is a small thing.

  • andai 5 hours ago
    Very cool project.

    > Backpropagation (gradients derived by hand)

    What does by hand mean in this context?

    Also how did you write the readme? It's a curious blend of human and AI writing.

    • porridgeraisin 3 hours ago
      The fact that by hand is emphasized so often and so often (see src/handgpt.h comments as well above backward()) makes me think it's AI.

      Anyways, it looks like gradients derived by hand means they didn't use autograd. They have written out the expression for dL/dW themselves.

      • wikisailor 1 hour ago
        Correct. No autograd: the expressions for the gradients are written out explicitly in C. There's a gradient check in tests/ that verifies them against centred finite differences on the published header, worst relative error 1.07e-08.
      • 4gotunameagain 2 hours ago
        I think the readme is clearly LLM generated. The tone, the short sentences, the em dashes, lists, sections.. It all feels like LLM.
    • wikisailor 1 hour ago
      [dead]
  • 12373bd 6 hours ago
    hIngan motlh puS ruq tuq DujDaq SISwI' nge'vI' vIn SuvwI' yIvwI' qarghtaHvIS SIchoH
    • rbanffy 14 minutes ago
      vIparHa'bej.
  • pavelai 8 hours ago
    This small Klingon speaking language model was trained completely on ESP32. The training took 2 days

    Number of parameters: 319K

    (Disclaimer) The models is tiny and is not a pocket chatbot. It mistakes and is not capable to support conversation, but that's not the goal of the project

    The goals of the project is to bring training to edge devices and it worked out

    How can one use it? By using solar panels such device could be turned into autonomous meteorological station