Mercury 2.5 LLM hits 770 tokens per second

(artificialanalysis.ai)

13 points | by Retro_Dev 1 hour ago

3 comments

  • bearjaws 14 minutes ago
    If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.

    I've used it on a few for fun projects and its decent but the speed is crazy to watch.

  • walrus01 32 minutes ago
    Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next or similar class of open weight LLMs that fit in under 170GB of RAM, so I don't see the point. I think this is probably also stupider than laguna s 2.1 which can also be very cheap to serve.
  • rvz 1 hour ago
    The speed means absolutely nothing when it is finishing almost dead last when compared to the frontier AI companies.
    • glouwbug 6 minutes ago
      Some of us want fast food
    • copperx 31 minutes ago
      Ah, the old "good, fast, or cheap; pick two" proves true once again.