One month coding with GLM 5.3 Flash

(wagtail.org)

30 points | by ThibWeb 4 hours ago

5 comments

  • epistasis 14 minutes ago
    One thing about these numbers that's absolutely shocking to me is how low the energy use is:

    > That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).

    The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's driving miles. It's boiling 10 gallons of water.

    With the talk of AI Data Center's impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.

    My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.

  • aktenlage 29 minutes ago
    > Unfortunately there are still consequences to it. I chose the 'wrong' model for the prototype, and we spent 450M tokens / $150 / 5kWh of energy use almost overnight. The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing:

    > Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.

    I don't get it. Why was it wrong? Which one would have been better? What was the lesson and how could you have foreseen it?

    • ThibWeb 18 minutes ago
      Hmm I might need to rephrase. Initial challenge was to use GLM 5.3 Flash and I was on the non-Flash version for the whole vibe coded build. Just wasn’t paying attention and I didn’t realize that one session was a quarter of the month’s spend and 150% of the budget (big price difference between models)
  • gpugreg 25 minutes ago
    How did you measure energy usage?

    Edit: I found a linked article that mentions the inference provider who does the measurements.

    • ThibWeb 17 minutes ago
      Yes, all from Neuralwatt, GPU energy use only. Makes models’ “efficiency” much more visible than tokens.
  • throw930rmdkdk 1 hour ago
    Flash is pretty decent coder, but it should be paired with good planner and reviewer. I would pick astra low for planning and sol 6.1 medium for reviews.
    • eikenberry 52 minutes ago
      What would you use if you wanted to stay (at least) open weight?
      • verdverm 35 minutes ago
        qwen3.8, kimi3, kimi2.7, GLM-5.3 are all good families I use in my coding team

        I'm mainly using flash varients, at least as the default, bump.up to stronger model as needed (less often these days)

        • Geof25 25 minutes ago
          qwen3.8 27B or 2.4T? they are completely different models with completely different pricing
          • verdverm 20 minutes ago
            I use all the qwen!

            It's my favorite model family to interact with, it's prose is the best imo, it makes me laugh from time-to-time (like when it said it would "crib" some code from another project, lul)

            I currently have qwen-flash working on an NES emulator harness so qwen-little can play my first RPG (ff1)

            (tho I have used all the others I mentioned, happenstance I'm using qwen this iteration/task)

      • esafak 37 minutes ago
        Deepseek 4.1 Flash and Mimo 2.6 Flash.
    • esafak 38 minutes ago
      Flash is plenty good for planning and reviewing, for my needs. In fact, I use it for that because it's too slow for execution, despite the name (from z.ai).
      • jminnl 25 minutes ago
        > I use it for that because it's too slow for execution

        What kind of hardware and what particular quant?

      • gunalx 29 minutes ago
        Will agree on this. From z.ai i have found the non flash to have way more consistent performance.
  • ThibWeb 4 hours ago
    It was a bit of a silly challenge, wasn’t sure how workable, learned a lot in the process about what actually drives usage / costs, and how to keep both under control
    • cvburgess 4 hours ago
      Considering they were your top two models, how did the flash and non-flash versions compare? Did you use them for different tasks?
      • ThibWeb 14 minutes ago
        I have a hard time justifying GLM 5.3 these days. It’s slightly better than Flash but rarely enough to justify the much steeper price. We chose to use usage-based billing only so are very sensitive to model price.
      • gunalx 30 minutes ago
        When text wuality or for pure but adwansed coding is concerned i always pick glm 5.3. The flash is awesome for everything that dosent really matter though.
    • sampullman 53 minutes ago
      Worth a comparison with DeepSeek v4.1 flash, if you've got another month to spare!
      • ThibWeb 10 minutes ago
        Yep, I think next month will be on that. It feels slightly better from a few days of use, and in our WIP benchmarking it scores way higher
      • samtheprogram 27 minutes ago
        DS v4.1 Flash is roughly equivalent. It's going to get some things right/better that GLM flash doesnt and vice versa.
      • lnenad 35 minutes ago
        They've got an image of their homegrown benchmark in the post that lists DS4.1.