Why your local LLM feels dumber than it is

(forum.level1techs.com)

36 points | by felineflock 3 hours ago

2 comments

  • jonplackett 1 hour ago
    I just got qwen 3.8 27b mlx running on my Macbook Pro and honestly I’m pretty blown away by how not-dumb it is.
    • prettyblocks 45 minutes ago
      My problem is how hot they run. I'm on an m4 pro. Do you have the same issue?
      • jonplackett 12 minutes ago
        It’s hot and also LOUD and runs the battery down quick.

        But I’m having a lot of luck just running things when I’m away from the computer and can leave it plugged in.

        It starts going weird (unreliable and slow) with context over 80k so you have to pick tasks one at a time and baby sit a lot more than Claude. But it really is very capable and feels like there’s an intelligence there to talk to. Maybe gpt-4 level clever?

        I have an m5 max 64gb and I think anything slower would be quite painful.

      • lukan 39 minutes ago
        I don't have the hardware but a often mentioned advice is to put your mac into energy saving mode - it still will work, a bit slower, but stays cool.
      • downrightmike 44 minutes ago
        Mineral oil bath?
        • kees99 9 minutes ago
          Or, set it up in another room, and connect from the other laptop.
    • alexchantavy 38 minutes ago
      How many tok/s are you getting? What gen mbp?
    • StarlaAtNight 44 minutes ago
      how quick does it respond? what are specs of your laptop?
      • Gareth321 13 minutes ago
        I tried it on my M1 MacBook Pro. It's slow but surprisingly smart as a general purpose LLM. Maybe GPT-5.3 level. I gave it a bunch of tools and it can search the internet, make product recommendations, document, code, etc.
      • chorlton2080 38 minutes ago
        Does it need to respond fast? For important applications, I'm sure we'd all be fine waiting 20 minutes for a high quality, usable answer. Or is it the need for interative refinements that make speed relevant?
        • jonplackett 10 minutes ago
          It requires patience but it’s more like waiting 5 mins for it to do tasks. You need to be much more involved though and do things slower than Claude where you can trust it to do a lot of tasks at once. It doesn’t have the context for that
    • applicative 31 minutes ago
      Did you read even the title?
  • anotherCodder 51 minutes ago
    [flagged]