1 comments

  • jasonjmcghee 54 minutes ago
    (2025)

    As it's 9 months old and they just had a major model release

    • throwa356262 18 minutes ago
      For K3 read this instead: https://arxiv.org/abs/2607.24653

      The main contribution of the K3 paper is Stable LatentMoE. Like some other models K3 compresses data sent between layers and this put certain requirements on the router. K3 improves performance by using a more balanced selection strategy here.

      • mcbuilder 2 minutes ago
        Compared to the Opus 5 "model card", which read like a standard Anthropic set of alignment principles and safety concerns, this presents a plethora of useful technical details that advances the state of the art.
    • cptcobalt 15 minutes ago
      Rather under-discussed back then: https://news.ycombinator.com/item?id=45766937
    • GaggiX 39 minutes ago
      I believe OP posted it because the new Kimi K3 has 69 KDA layers (the rest are 24 Gated MLA), I think previous large Kimi models had only MLA layers.