87 comments

  • a11r 1 day ago
    I'm a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I'm running 4-bit quants on an RTX Pro 6000 rented for approximately $1/hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: https://www.reddit.com/r/BlackwellPerformance/s/FrKwk3GoDK
    • seamossfet 1 hour ago
      >I'm a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality

      There's literature on this actually, if you want to get that low the model really should have quantization as a pretraining target otherwise larger models that are quantized after collapse.

      In fact, to do this effectively, you need somewhere near 50x chinchilla to do QAT on super tiny targets like 1-bit or even 2-bit. These large models already need an astronomical amount of training data that doing proper QAT that small isn't really feasible unless there are breakthroughs in the architecture.

      4-bit models and smaller _can_ perform well but not through just naively quantizing the existing weights of a large model.

      • happyPersonR 21 minutes ago
        My we should crowd fund distillation with 2 bit quants on a deepseek 4.1 flash and see how it does ? Bet it does well ..
      • alfiedotwtf 1 hour ago
        > 4-bit models and smaller _can_ perform well but not through just naively quantizing the existing weights of a large model.

        Doesn’t that describe the majority of quants on higgingface?

        • seamossfet 1 hour ago
          Yeah, there's a tradeoff. pre-training with QAT needs more data and a ton of hardware, so it's easier to just take an existing open weight model and quantize it. This works, but it'll underperform a model that had quantization as a training target.

          Matters more for 1-bit and 2-bit models.

    • Winfred-zz 19 hours ago
      I just ran a set of benchmarks, ninfer-3090-qwen3.8-27b (so mix of Q4 and Q5) vs strata-qwen3.8-flash-next-iq3_xxs (so Q3):

      │------------------- │ Ninfer-3090 │ Strata

      │ Code generation │ 52/78 (66.7%) │ 70/78 (89.7%)

      │ Code completion │ 40/50 (80.0%) │ 44/50 (88.0%)

      │ Total------------- │ 92/128 (71.9%) │ 114/128 (89.1%)

      │ API failures------ │ 10 │ 5

      - Ninfer generation: ~122 min total.

      - Strata generation: ~142 min total.

      So strata is a little slower, but keep in mind that ninfer-3090 is very optimized for a Qwen 3.8. Standard Qwen 3.8 runs at 20 t/s, this modified version can do 50 t/s (but it's extremely long in it's thinking, it just goes on and on.

      This is on a 3090 that will crash unless power capped, with a Zen 2 CPU, 64GB DDR4 with a PCIe that refuses to go higher than 8x (basically pretty crappy all in all).

      Yet with some tweaking and optimizing I still manage to get strata to run at 40 to 60 t/s.

      That strata has been optimized on my Oh My Pi conversations. So when I'm using it, it's probably faster and closer to ninfer in speed than during those unoptimized benchmark tests.

      • tharkun__ 15 hours ago
        You can use 4 spaces on HN to get code/monospaced formatting:

            │------------------- │ Ninfer-3090    │ Strata    
            │ Code generation    │ 52/78  (66.7%) │ 70/78   (89.7%)
            │ Code completion    │ 40/50  (80.0%) │ 44/50   (88.0%)
            │ Total------------- │ 92/128 (71.9%) │ 114/128 (89.1%)
            │ API failures------ │ 10             │ 5
      • seviu 11 hours ago
        Unrelated... just use your just locally installed qwen 3.8 and instruct it to figure out why you arent at 16x.

        For me, I didnt know I was on 4x and I was able to, giving it enough permissions, with a good harness like PI, to figure things out and go 4x -> 8x -> 16x.

        It probably will take some opening the case and switching ssds around, but it's amazing how powerful these models have become.

      • airspresso 11 hours ago
        > it's extremely long in it's thinking, it just goes on and on

        Known issue with this model, I recommend setting thinking to 'medium' instead of default 'xhigh'.

        • spockz 6 hours ago
          Ill check for this in my setup. I thought it came with only one thinking level causing the “meandering” as it refers to its own long CoT.

          I’m doubting the quality of the model, even comparing with luna 5.6, it is giving such bracingly wrong answers. It claims other tools have changed the code when pressed why it made a mistake. That code or tool didn’t come close to working on the project. It is scary.

        • sethd 4 hours ago
          You can also set a thinking budget in terms of how many tokens it can spend inside the reasoning/thinking phase.
      • greenavocado 14 hours ago
        In my private C coding benchmark, Strata's Qwen Flash Next Coder IQ1_M beats Luna low and runs at 125 tok/s on my 5090
        • seviu 8 hours ago
          Not as fast for me. I was able to snatch a spark before the horrible price hike, and I have it at 40/50tok/sec. Before I had the great 3090 setup which is what led me to get the spark. That 3090 is now relegated to creating videos of my kids doing dumb things.

          I now backordered two more sparks, before the hike... hoping they will honor the agreed price and dont cancel on me. They should arrive in a month,

          Its evidently clear the frontier labs won't keep on giving us cheap inference for much longer. I also watch in disbelief how people say a sub is cheaper, which probably is. But you loose on so many other things (privacy, predictability, censorship...)

          • greenavocado 3 hours ago
            I pray for the flood of CXMT high speed memory and Huawei accelerator chips
      • sleight42 18 hours ago
        Anecdotally, I've found that 27B at 4bit hallucinates a lot more than QFN at 3_xxs, particularly for less technical reasoning.

        I've tried to use both for online comparison shopping. QFN not only seemed less delusional but also made useful observations and problem solved ways around many different website access issues.

    • sudo_cowsay 16 hours ago
      For the 1 dollar per hour, may I ask (without trying to incite things) why you would not just buy a ChatGPT Plus subscription or other 20 dollar subscriptions? Do you like the privacy?
      • bad_haircut72 9 hours ago
        People think claude costs $200/month, it actually costs $200/month and all your labor to train their models, improve their product, and all your best ideas and work that they will feed back into their company. Its very expensive
        • komatar 7 hours ago
          You can opt out of model training with a checkbox, though I doubt most companies respect this completely. At the very least, I imagine they find loopholes to exploit.

          But on the other hand I'm giving them my best and worst ideas and the value I get in return outweighs what I'm contributing to them. I also don't have the problems of initial HW cost and maintenance.

          • mixermachine 3 hours ago
            Yer I also use service with zero data retention promises but it is hard to really trust them for me... They trained on basically all of the existing material, legal or illegal already and are starved for true human input. They are also in a heavy arms race were basically only a handful of companies will survive.
          • zdragnar 2 hours ago
            Their own TOS will still include your "anonymized" data in their "research", with no explanation as to what that means. You need an enterprise account to get real zero data retention.
          • wren6991 6 hours ago
            Great, then they will PII-scrub my sessions before feeding them into the training pool :-)
          • radicalbyte 6 hours ago
            > You can opt out of model training with a checkbox

            Do you trust that a company who can't even do basic IT ops monitoring are capable of respecting a "no training checkbox"?

            • cryptonym 3 hours ago
              "Whoops my bad, you are right, your data may have been in the training pool. They are now forever in weights that'll be reused to train every future model. Too bad you can't do anything about it."
      • rnd0 9 hours ago
        Access to models is currently being priced below it's actual value. Largely because the market hasn't settled yet. Once it does the prices to access AI will increase -signifigantly.

        It will be less of an issue for people who have already got the equipment and the practice running locally than it will be for people who have simply been relying on ChatGPT.

        • jeremyjh 6 hours ago
          Prices on openrouter are not subsidized and are very affordable.
      • manmal 11 hours ago
        Those plans are super slow right now. Nowhere near 100t/s.
        • azath92 8 hours ago
          While this may be true, their throughput is practically unbounded. this has been a primary challenge to using local models at work, where its normal for me to have 1-6 sessions churning away at once. So while 100t/s on a single session is very fast, theres no way you could get a cumulative ~500t/s from a consumer setup like this, do the bandwitdth bottleneck on a single card/cpu/ram/ssd (tends to be gpu ram bandwidth is the limiter in my setup, but they all have respective limiters to this kind of throughput)
        • martianvoid 10 hours ago
          I think Opus 5.5 is now fast enough with the inference stack improvements that they did. Sonnet 5.5 is around 140 tokens/s when I had measured it last time. A $20 subs on claude and using only sonnet 5.5 will last a lot
          • c16 9 hours ago
            I was firmly wanting to move towards local llm, and since the Sonnet 5.5 and recent Opus speed increases, I'm putting that on hold until local inference speeds can be improved. Not sure how much juice there is to squeeze there, but I'm hopeful.
        • Almondsetat 8 hours ago
          Does it matter if those 100 tokens per second are way crappier?
      • mannanj 15 hours ago
        maybe ethics as well. maybe people like sleeping well at night.
    • walrus01 4 hours ago
      I have to concur with you here because I'm running 3.8-flash-next in the largest unsloth Q8 version on a 256GB system for daily use, and I've recently compared it head to head vs a Q4 that I have on a 128GB system and the Q4 is just plain not as good. The Q4 is given only more basic tasks.

      I can't even imaging letting a Q2 anywhere near a code base that I care about. To generate a bird on a bicycle as a test? Sure. Actual things I don't want it to screw up? nooooooooo.

      • bitexploder 3 hours ago
        There is a magical quant GSQ-RCO that at 3-bit XS beats the full unquant. I have tested it. It is on par with the full unquant on DeepSWE. Give it a shot, it is magical. ITSA lab, they basically devised a way to selectively quantize parameters. I am a GSQ-RCO truther at this point... it is so fast on my V100s and I can't find an instance where it has acted weird or fell short of the full quant on agentic coding. I Have it 3d model in blender, all sorts of stuff. It's just... good.
    • NinjaTrance 22 hours ago
      Just as curiosity, how long does it take to set the environment up and running?

      Is it viable to start/stop it multiple times per day?

      • a11r 21 hours ago
        Yes, you can probably get the whole thing up and running in about an hour the first time. If you pause and restart, it takes about 15 minutes to load the models from disk into GPU memory, so budget for cold startup time.
        • KeplerBoy 21 hours ago
          How does it take fifteen minutes to read <100 GB into GPU memory? Shouldn't that be limited by SSD speed with everything slower than a minute being a terrible ssd?
          • boredatoms 19 hours ago
            It also depends on the runtime, vllm is unbelievably slow at model loading compared to llama.cpp
            • ycui7 12 hours ago
              this is not right. vllm can easily load safetensors model at >5GB/s if not faster when setup right. did you make the compile cache persist? if you use docker, you should bind mount the kernel compile cache, so they don't need to be recompiled each time vllm restart.
          • teaearlgraycold 21 hours ago
            A lot of cloud platforms have terrible slow network storage. They also might need to compile the GPU kernels fresh as they might not have a persistent CUDA cache.
      • prettyblocks 13 hours ago
        I just had claude code set it up for me and create some startup scripts. The longest part was the model download.
    • robot_jesus 1 day ago
      Can you say more about where you're renting the RTX Pro 6000 for $1/hour?
      • a11r 21 hours ago
        I am renting spot VMs from Nebius. I've tried a variety of other providers like Vast and Spheron. Vast worked well for renting 5090s but I like the large memory and pricing I'm getting at Nebius for RTX Pro 6000. The extra RAM really matters because I need PLE to offload the ngram to RAM.
      • _zoltan_ 23 hours ago
        vast.ai?
    • nialv7 20 hours ago
      • ricardobeat 20 hours ago
        What are you running it on? I'm getting mixed results using the IQ3_XXS quant which supposedly matches baseline, it feels significantly degraded.
    • sfifs 18 hours ago
      I've run DeepSeek V4 Flash on DS4 on single DGX and standard model weights on aDGX cluster. There was some degradation going to the hybrid 2 but quant but really not much.
    • segmondy 18 hours ago
      The larger the model, the more you can go down. K3 in Q1 will match and likely beat Qwen3.8-Flash-Next.
    • sail0rm00n 1 day ago
      $1/hr sounds great. Where are you getting it for those prices?
      • transcriptase 23 hours ago
        [flagged]
        • dang 21 hours ago
          Could you please stop posting unsubstantive comments and flamebait? You've unfortunately been doing it repeatedly. It's not what this site is for, and destroys what it is for.

          If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.

          • transcriptase 20 hours ago
            Apologies. I had recently read (I think from someone on here) that they were in some discord where those sellers advertise. Whoever it was mentioned that the service was incredible value for what they got, but admittedly it was for toy projects, there was no written agreements, and uptime was something like 98%. I should have found the source and linked, and will try to do better overall in my replies!
        • xingped 23 hours ago
          No? Per someone else's comment, presumably they're using vast.ai which does indeed list the 6000 for $1/hr
    • aatd86 23 hours ago
      $1/hour ? I need that deal as well.
    • nickpsecurity 20 hours ago
      That's cheap, too. Which hosting service are you using?
    • alvarolucero 10 hours ago
      [flagged]
  • SomeHacker44 4 minutes ago
    I have 2x R9700 and 192G RAM on a computer and would like to run this!
  • Jackson__ 22 hours ago
    I've just tested Strata on a simple 50 image vision benchmark. The task is to output the exact coordinates of a requested object. The result via Strata had a median error distance of 154.8 pixels, avg of 168.8. Running the exact same GGUF and vision adapter weights on llama.cpp gives me a median error of 46.5, avg 81.4.

    To put that into perspective, here are some more numbers from other models via llama.cpp:

    Median/Average

    Qwen 3.5 9B BF16: 46.5 / 193.3

    Qwen 3.6 35B Q4 K XL: 38.4 / 76.4

    Qwen 3.5 122B Q3 K M: 32.9 / 68.6

    The difference in vision performance is as large as the jump from a 9B model to a 35B model. All tests were performed at temp=0.

    I have done no further testing, as these results line up perfectly with my expectations.

    • amitp85 3 hours ago
      I had similar experience that same model running on Strata made more mistakes and hallucinated more when compare to llama.cpp. Once the model swift-Qwen3.8 FN iq-xxs, got confused that it has already created few py files and they disappeared from disk, then I check all logs and it never created those files.
    • bitexploder 3 hours ago
      Real world: I use QFN38 to build 3d models. I use build123d + blender and this model is the first local model that does a good job. It is kind of amazing. This model is as good as DSV 4.1 on most tasks and I think it only loses when specific knowledge matters.
    • larodi 11 hours ago
      To be done right, such vision benchmark should consider the peculiarities of the vision tower, and was it quantized and optimized. So a lot of success/loss of quality may not be due to the transformer.

      Second of all, the "find coordinates" of something is super difficult task of any model, so you tried to test a small quantized buddy with a tri-star challenge. Not sure what expectations were set.

      disclaimer: myself do large volume VLM work daily, including in production, for more than 1.5 years now.

      • Jackson__ 9 hours ago
        >To be done right, such vision benchmark should consider the peculiarities of the vision tower

        As per my previous comment, I used the _exact same weights_ for the comparison. The vision tower is always kept as an unquantized BF16 file for GGUF, as I believe is the default.

        >Second of all, the "find coordinates" of something is super difficult task of any model

        It is in fact not (anymore), and most recent VLMs I tested have been standardized to point and bbox in a relative 0..1000 coordinate system with decent enough accuracy.

        And finally, none of this excuses that Strata performs so much worse with the same weights. Which does kinda leave me confused as to what the point of this reply is.

        • larodi 7 hours ago
          > Which does kinda leave me confused as to what the point of this reply is.

          to point out that 1)measuring VLM is not straight-forward, and needs considering what was actually ablated.

          and that 2) perhaps such metric is the most challenging one for a brutally quantized and otherwise lobotomized model.... which otherwise performs well in other tasks (which I also doubt - for the record!).

          i can subscribe to the idea that the Strata performs worse because it stripped the model of important quality, but not entirely to the fact it was really properly evaluated.

          the way such models/attempts degrade can_be/is insightful on its own.

      • knollimar 9 hours ago
        What type of work do you do? Do you consider shelling out to a tool where they can place markers and retry?

        Do you have any other benches? What models do you use or recommend for this type of work?

    • Xenograph 18 hours ago
      Is this test available somewhere? Would like to test it out on my models.
    • biztos 13 hours ago
      How big are the images? Is 160px 1%, 10%…?
    • throwaway219450 19 hours ago
      What does SAM3 get on the same test set?
    • NamlchakKhandro 20 hours ago
      Tldr, strata is a waste of time.
      • unlikelytomato 16 hours ago
        at least for vision? Are there similar comparisons for language? It seems like vision is often an afterthought when it comes to bootstrapping these newer inference engines
        • Borealid 16 hours ago
          Vision functions the same way as language when inference is done. It's a stream of tokens.

          It's just easier to measure the "right" answer (and deviation therefrom) on a vision task than a language one due to the underspecified nature of language.

          • unlikelytomato 15 hours ago
            my understanding is the vision component of the original model is an independent preprocessing step(except for things like Gemma 12b). And it would be possible for engine to have broken that phase of translation before it hits the real model as encoded text. I am not familiar with Strata engine and their claims, but this seems like an interesting way to test models with relatively straightforward inputs.
            • dr_kiszonka 10 hours ago
              Unsloth documented the vision components degrading more during quantization for one of the earlier Qwen models.
          • pjc50 10 hours ago
            How does image tokenizing work?
  • daft_pink 7 minutes ago
    Will this strategy work on lower memory unified memory machines like Apple Silicon or AMD?
  • snehesht 1 day ago
    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    https://huggingface.co/Qwen/Qwen3.8-Flash-Next

    • roscas 1 day ago
      Coder version with 30t/sec on a Ryzen 3600x with 48GB of RAM with a nvidia 3080.

      This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.

      I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.

      I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.

      This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.

      • StumpChunkman 1 day ago
        How much VRAM on your 3080? I've got an early 10gb model. I've been thinking of exploring local coding models, but everyone seems to use much better GPUs than I have access to. Yours is one of the first I've seen with maybe similar hardware on some level.
        • Abishek_Muthian 9 hours ago
          With limited VRAM, you can still use local LLMs but you have to set your expectations correctly to solve the right problem.

          I run small models on various kind of devices including 1.7B model on the original Jetson Nano (4B) abandoned by Nvidia, I had to upscale the software (OS, Lllama.cpp etc.) to run the model but the model runs at 17t/s.

          Small models are great at NLP stuff like classification (e.g. bookmarking), ASR etc. I have built custom browser extensions to save time with bookmarking and categorizing to use with local LLMs.

        • roscas 1 day ago
          Yes, 3080 with 10GB, forgot to mention that.

          Mine is at the moment writting some cpp code for some SBOM tests.

          I have loads of terminals open. Librewolf, Chromium and you know how this crap likes ram, I have also a vm with 4gb of ram running and doing stuff while I wait for the results but hey, while I wrote this the program is done. Wow! That was 29.x tokens per second most of the time.

          Oh I will run some other tests with hermes now because hermes is amazing too.

      • la_oveja 10 hours ago
        so it runs on 10gb vram?
    • thatsabadlook 1 day ago
      Why is this surprisingly well? It's 2.5x faster than anthropic models, you have data sovereignty, privacy,and that's a strong model. Sounds like a best case scenario to me
      • hdjrudni 1 day ago
        Not sure you understand the term 'surprisingly well'. It means 'better than expected'. I suspect they parent poster didn't actually expect to get >= 100 T/s.
    • eek2121 6 hours ago
      I've 32gb of DDR4 currently, along with a 4090. I am hoping to try it later. My 32gb was previously preventing me from using this model.

      I'm hoping it will be better than Qwen 3.8 27b, which is already quite good.

    • bobkb 3 hours ago
      How long does it take to load ?

      On my RTX 5000 Pro + 128GB RAM machine its taking 10+ minutes to load.

    • proc0 1 day ago
      Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
      • incognito124 1 day ago
        Qwen 3.8 flash next is way better than 27B. It's so good I dont even use claude anymore
        • mickeyp 1 day ago
          I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.

          It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.

        • roscas 23 hours ago
          I prefer https://ornith.ai/ornith_1_5.html to Qwen 3.8 not only because it is much faster on my hardware but better responses.

          But this Qwen 3.8 Flash next coder is amazing running with Strata.

        • JokerDan 1 day ago
          Is this true for 27b Q4_K_XL vs flash next IQ3_S? I thought under Q4 models start quickly degrading?
          • gruturo 1 day ago
            While this is generally true, it's _a little_ less true the larger the model is.

            Also, quantization techniques have improved - the I in IQ3 stands for imatrix - Importance Matrix - it is a bit more surgical in what it cuts. The result is a model where the most important weights are even Q6 or above, the least important Q2 or even below, overall it takes the space of a Q3 but with better results.

            • ranguna 9 hours ago
              A lot of words, but no benchmarks. I'm a little tired of all the "finger in the air" vibe checks. You can say all you want, but you'll only know once you put out some numbers. Which people have in other threads, and 4 bit 27B beats Flash at 3 bit
          • latentsea 15 hours ago
            This is the conventional wisdom, but in practice what matters is how reliably the model performs on your tasks in the real world. I have an R9700 and an RTX 5060 Ti and I've been running an IQ3_S quant of 27B on the 5060 Ti vs a Q6 quant on the R9700. I still manage to get stuff done with the IQ3_S quant.
          • Tade0 21 hours ago
            To add to the other comment, there's also Ridge quantisation - the majority of weights are indeed Q3_x, but the most sensitive layers are FP8.
        • snehesht 1 day ago
          Yeah I agree, I'm running it with Pi didn't notice much difference compared to lower tier models and the speed, of course.
          • nicce 1 day ago
            I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.
            • DoctorOetker 23 hours ago
              do LLMs tend to be homesick when not used in the same harness they sat in during some training phase?
              • Bnjoroge 20 hours ago
                iirc there was a sectionin Qwen’s paper where they talked anout how they post-trained flash or 3.8 to work just as well regardless of the harness or eval used. I think that used to be true but not sure if it is any longer
      • a11r 1 day ago
        We recently moved from 27B to Flash Next. The quality is superior for coding. Our workload is primarily well-defined coding tasks that need to be attempted a few times before the model gets it just right. FlashNext is also better at finding issues in generated code than Gemini 3.8 Flash.
        • swozey 20 hours ago
          I'm on m1 max 64gb and went from qwen3.8-27B back to qwen3.6-a35b. Is flash next the move? I went from usable say 40tk/s qwen3.6 to unusable, like 11 with 3.8 and not impressed with the replies for the time sacrifice. pi (omp) and omlx but not with the recent 3.8 patch.

          I've been waiting for a 35b of 3.8, I don't really know what the other versions are about. I'm on 5g so juggling 40gb of model files sucks. And honestly I'm sick of tweaking this stuff for no, very little, or break-it level improvements. Qwen3.6-a35b has been solid for work, just don't give it freedom to wipe your data.

          • ENGNR 18 hours ago
            Exact same scenario here

            I’ve heard a quantised version of flash next can fit in ~50 gb of vram (which needs a system level flag set to go over 48gb)

            But the m1 cpu is itself a bottleneck on prefill compared to say an m5, there’s no real getting around it. And the 400mb/s bandwidth starts to hurt without MOE

            Hoping these model optimisations can see us through to 2028 because for everything other than LLMs this hardware is still over specced and working incredibly well

      • thatsabadlook 1 day ago
        Significantly better for both performance and real world use case. 3.8 27b is a good small model. This is a good model.
        • geye1234 1 day ago
          I find 27B more accurate -- maybe because I'm running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.
          • PcChip 1 day ago
            Spelling mistakes?

            What inference engine are you using for flash next?

            • anon373839 1 day ago
              Yep, can confirm that is NOT normal. Are you using Nvidia’s NVFP4 quant? There are other NVFP4s floating around but they are not as good. The quality of the calibration data really matters.

              Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)

            • geye1234 19 hours ago
              I'm running Pennyroyal's Docker image (on Podman) which uses sglang. I have a single RTX 6000 Blackwell and 128GB RAM. I turned off disk caching. I'm running with a ~500K context, but have been limiting it to 256K in the client (pi).

              It always detects its spelling mistakes, btw, but it worried me. It may turn 'rm -rf ' into 'rm -rf /' one day.

              Almost certainly the problem is my config, not the image.

              • phacker007 3 hours ago
                are you using MTP with concurrent requests? I think there could be an issue with KV cache I had with llama.cpp that randomly threw in characters from other thread's KV cache...
              • gewetensleegte 6 hours ago
                > It may turn 'rm -rf ' into 'rm -rf /' one day.

                but surely you're running it in a container?

              • xiconfjs 10 hours ago
                Did you check your GPU for memory errors/defects?
          • thatsabadlook 20 hours ago
            Definitely not.
    • jacquesm 19 hours ago
      Speed is one thing, accuracy another. Have you benchmarked it against a reference? If so, what were the results? I tend to go for accuracy over speed because usually that means fewer round trips and fewer tokens wasted.
    • notnullorvoid 1 day ago
      Which quantization are you using to reach those numbers?
  • Abishek_Muthian 39 minutes ago
    I tried "write binary tree in rust and ensure it compiles with rustc" using the coder model with their chat it didn't produce the correct output and their chat can't call tools. Then with opencode it went on thinking loop until it exhausted the thinking budget.

    4090 16gb, 96 gb RAM.

    But I like this overall approach of neutering experts and loading them from ssd to make the model run in low VRAM systems.

    • atomicnumber3 18 minutes ago
      IME, local models (so far) cannot do anything resembling "vibe coding" or what I call "high ambiguity prompting."

      Ex:

      Won't be good at it: "make a web app where you can download files with background jobs"

      Much better at: "make a new dbmate migration to create a table with this schema, then write REST-style CRUD methods for it: [pseudocode table definition, you can gesture at types and indexes]"

      So, they don't demo oneshots as well, but for actual "LLM-augmented software engineering tasks" they do super well.

  • cjdell 13 hours ago
    This is game changing. My R9700 32GB is now smarter and about 2x faster than using Qwen-3.8-27B. About 60 t/s when combined with my 96GB of DDR4. My motherboard limits me to PCIe Gen3 so that is likely a bottleneck.

    For the Nix inclined: https://github.com/cjdell/nixos-config/blob/main/hosts/zen3-...

    Even got it running on the iGPU of a GMKTec M6 Ryzen 6600H at reasonable speed (10 t/s). Fast enough to leave it with a prompt before I go to bed and wake up to a solution.

  • AntiRush 20 hours ago
    I've been working on support for this model in ds4 on the RTX 6000 pro - it's been really great for my use cases. The ds4 q4 quant performs a lot better than other similar sizes that I've seen.

    Using the Q4 quant on an RTX 6000 Pro Workstation Edition at 450 watts:

      Code: prefill 1,251 tok/s decode 255.26 tok/s
      Prose: prefill 1,251 tok/s decode 198.78 tok/s 
    
    Most important for me, I can run 4 concurrent streams at 400+ tok/s.

    https://github.com/fairfieldt/ds4

    • jacquesm 19 hours ago
      DS4 is an odd model. I have it working on way too many GPUs and yet for many tasks Qwen 3.8 will do much better. It also tends to loop, which is super annoying.

      GLM5.3 runs on similar hardware and is much better so if you're going to burn cycles and brain power on this maybe look at GLM5.3 as a comparison as well?

      Other than that, when you're done with that card...

      • AntiRush 16 hours ago
        Agreed that glm-5.3-flash is a strong model. On a single 6000 pro you can (barely) fit a q2 quant in vram. I haven't used it extensively, but my initial feeling is that the q2 is quite a bit worse than q4 for this model. To run it comfortably at q4 with reasonable context length you really need 2 6000s.

        When the qwen 4 series is released I am hopeful there'll be a strong model with the same architecutre. as qwen3.8-flash-next.

        • jacquesm 11 hours ago

            |  0  N/A  N/A  321801 C llama-server 16540MiB 
            |  1  N/A  N/A  321801 C llama-server 22404MiB 
            |  2  N/A  N/A  321801 C llama-server 23480MiB 
            |  4  N/A  N/A  321801 C llama-server 43982MiB 
            |  5  N/A  N/A  321801 C llama-server 41656MiB 
            |  6  N/A  N/A  321801 C llama-server 22366MiB 
            |  7  N/A  N/A  321801 C llama-server 16208MiB
          
          That's 3.86 bits per word I could run the 4 bpw one as well (there is still one spare gpu and another 27G on the ones listed above. DS4 uses about the same memory, is a little bit faster (though I suspect that by the time GLM 5.3 support is a bit more mature the speed difference will have evaporated).
        • alfiedotwtf 58 minutes ago
          GLM 5.3 Flash even at Q3 really feels like “we have Claude at home”
      • segmondy 18 hours ago
        ds4 is the inference engine not the model.
        • jacquesm 11 hours ago
          Sorry, but no, I use DS4 as the shorthand for Deepseek 4, it is no coincidence that that particular inference engine was called DS4.

          From the homepage of the inference engine:

          "DwarfStar aims to be the best way to run a few excellent large language models on consumer hardware (that is, hardware that people can actually own). To reach this goal, we are building a small native inference engine optimized first for DeepSeek V4 Flash (including the experimental vision model), DeepSeek V4.1 Flash (Metal, and text inference on CUDA), and additionally GLM 5.2 and 5.3, GLM 5.3 Flash and DeepSeek V4 PRO, and Qwen3.8 Flash Next (Metal and CUDA)."

          The model has been around longer than that. DS3 = Deepseek 3 etc. I think Salvatore named it pretty cleverly but he doesn't automatically get to own a two letter acronym.

    • anon373839 11 hours ago
      Those prefill numbers don’t seem right for Qwen Flash Next. I get 3,000 tok/sec on a DGX Spark and that’s got a lot less raw compute than the 6000 Pro.
  • kamranjon 1 day ago
    Dwarfstar already supports this, curious how it compares, but I use the q4 quant daily and it works really well.

    https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...

  • jacquesm 19 hours ago
    LLM threads the world over are spammed with Strata links, it remains to be seen how much of the breathless hype remains standing once the honeymoon period is over. I've tried it but so far I have not seen anything that overly impressed me in terms of accuracy, though the speed is definitely there. I'm sure there are applications for LLMs where the quality of the answers is less important but I don't have any of those. YMMV.
    • kristopolous 10 hours ago
      I'm sick of these 4 day old vibe coded projects becoming hyped to hell while I spend a year on something and it goes nowhere

      I don't know who this kid is but I've looked at the code, it's all Claude. Saying you can run on 2-bit quant with some optimization flags at 100t/s isn't like some "oh my goodness" ... it's wasting everyone's time with noise. I mean give me a break.

      What are they doing? Is it Tiktok? Discord? LinkedIn?

      I want to do high quality work but apparently I should be dicking around on social media

      • jacquesm 10 hours ago
        There are quite a few of these, indeed. Most of them either start of with llama.cpp (everybody's favorite to rip off, make a minor improvement to and try to establish a reputation) or vllm. But some of these actually do have something to bring to the table, for instance, by limiting the themselves to be able to focus on a smaller set of possibilities and this can lead to favorable outcomes for a smaller size of the audience. One particularly good example of this I think is 'ninfer' which mainly focuses on the Qwen family of models served up on RTX 5090's, which it does outstandingly well. Other people then go and fork that to adapt to their hardware so now there are 3090 and 4090 forks of ninfer.

        Then there is the 'unified memory' branch of inference engines, the most notable of which is probably DwarfStar 4 by 'Antirez', which also started off with a lot of code from the llama.cpp codebase.

        And then there is 'the rest', but even there, some of these have interesting bits and I always hope that eventually those bits will make their way back to the engine where it started.

        I run both ninfer and llama.cpp, vllm is an unmaintainable mess even though there usually is a performance edge (it is great if you are serving up for commercial purposes so you can tweak it for one set of hardware and one particular model). You'd essentially need to dedicate a week or more to getting a new model up and running on a particular set of hardware if it does not nicely match with the recipes found online.

        One exception is the DGX Spark series, there vllm is supported by the manufacturer and the hardware is very consistent from one box to another. But the performance isn't really there when compared to a fat PC with a bunch of GPUs. (The 200 GB/s memory bandwidth is a serious performance killer).

      • imtringued 5 hours ago
        "I'm sick of these 4 day old vibe coded projects becoming hyped to hell while I spend a year on something and it goes nowhere"

        You know, in a way, you are both in the same boat. The difference is that the 4 day old project gets the clout first and then becomes irrelevant.

        • alfiedotwtf 52 minutes ago
          Same could be said about models… every now and again there’s hype around a new model, then you look at the card and it’s a 27B or 35B-A3B without a single mention of Qwen in the text.
    • ActorNightly 9 hours ago
      Qwen models famously chase benchmarks. They are good at very specific tasks, but they aren't good problem solvers. The reason Strata works is because of those MoE models, which basically shoehorn themselves into less intelligent versions.

      Meanwhile models like Gemma4 which are fully activated preform slightly worse on the specific benchmark but are capable of exploring much more domain space with careful prompting.

      If you have a card with 24gb of ram, you can run Gemma4:31b above 100 tok/sec if its not being used for rendering. With careful prompting and RAG system, you can actually make it do better than frontier models. Especially if you add web search into equation (which with google basically runs Gemini for you).

    • brcmthrowaway 15 hours ago
      What does this inference engine do that others don't?
      • jacquesm 11 hours ago
        It is faster than comparable engines using the same model, but they use a lot of short-cuts.
  • SuperV1234 1 day ago
    We're getting closer and closer to the day we can have an Opus-like model running locally. The dream!
    • stymaar 1 day ago
      It depends on what you mean by “Opus-like”, because if you mean “as strong as Opus 4.6 for agentic coding” then Qwen3.8-27B has been there for the past two months.

      But if you mean “as strong as current-gen Opus” then it's probably never gonna happen, but it doesn't really matter since we're long into the diminishing returns for performance improvements: I haven't notice any major leap between 4.6 and 5.5 in my daily usage, and I'm convinced that with a fact enough piecs of hardware I would be using local Qwen exclusively (I'm using it daily but only at night for long running tasks because they take much more time than Opus due to the compounding effects of my slow GPU and Qwen's verbosity).

      • nycdatasci 1 day ago
        Opus 5.5 was a step function change. Really crushes on multi-hour coding compared with prior models.
        • stymaar 11 hours ago
          I feel like Anthropic is going back and forth on this, as Opus 5 was the laziest since 4.6 by far, often giving up after ten or fifteen minutes and being like “of course there's still this, and this and this to do, but it's long and you should take time to think about the schedule for these tasks”. It gave up way earlier than Qwen3.8-27B for instance.

          Opus 5.5 just fell like a return to the mean afterwards (it's probably an improvement over the previous versions, but I couldn't really feel it because I've been burned by Opus 5).

        • jjcm 23 hours ago
          Highly agree.

          I keep saying “I’d be so Happy with ${currentOpusVersion} locally”, but I keep being impressed with how much the capabilities change between versions. I have a RTX 6000 pro so I can easily run this qwen 3.8 flash next, but it’s much harder to give up the freedom that 5.5 gives me.

      • hgoel 1 day ago
        The most visible leap between 4.6 and 5.5 seems to be that the latter has gotten much more computer-use training, so there's a clear progression in the ability for the model to use Blender. But catching up on that is just a matter of training on the same thing.
      • system2 1 day ago
        I would be forever happy with Opus 4.8.
        • bitexploder 1 day ago
          Flash next is /really/ close. It is at parity with 4.7 as far as I can tell and basically where Opus 4.8 was. It is a genuinely good model. And I run it at home on $1500 of GPU at 125 t/s :)
          • gruturo 1 day ago
            Seconding this. Flash-next (and let's not forget, it's a PREVIEW of the 4 architecture - with the "real" 4 rumored coming later this month) is the first model I can run on reasonable hardware (2 thoroughly obsolete P100s off ebay at ~$100 each plus the RAM I could scavenge from other PCs at home) at a reasonable speed (22, with GPUs in layer-split due to llama-cpp's limitation on qwen4-exp arch, and no MTP. Strata could double these numbers).

            It's.... the real thing, for the first time. If you cut me off cloud models today, I would get plenty of utility out of this thing.

            (Others may have had the same feeling from GLM5.3 or Deepseek 4.1 flash but I never had a chance of running those.)

          • fsiefken 21 hours ago
            Running an nvidia card at full load, would cost me ~100 euro of electricty each month (europe). Of course one wouldn't have usage caps.

            Where the internet was a subscription 15 euro subscription to encyclopaedic knowledge, an genAI subscription is renting a researcher/programmer for 100 euro.

            • stymaar 21 hours ago
              > (europe)

              Given the massive difference in electricity price between different european countries, adding "Europe" doesn't bring much context.

              • magicalhippo 20 hours ago
                Even with a country. Here in Norway we have multiple price zones, and at times there can be 100x difference between them, often 10x. All due to lack of transmission capacity between northern and southern zones.
                • stymaar 11 hours ago
                  TIL, thanks.

                  In France we have the same price in all continental France, there are whole regions that are poorly connected to the grid (Britanny and the French Riviera) or even remote islands like Corsica or even the islands in the Indian Ocean and the West Indies but the price is still the same as elsewhere.

                  • magicalhippo 8 hours ago
                    There has been a lot of talk about abolishing it, and have one zone. But the grid will still be bottlenecked between certain regions, at least for decades, so opponents argue it could lead to grid instability or inefficiencies.
            • bitexploder 18 hours ago
              These GPUs use about 400W at full tilt (200W each), for reference.
    • latentsea 1 day ago
      We are already there. Prior to Strata the best I could run was Qwen3.8-27B at Q6, which itself is already at like Opus 4.5/4.6 level, and now with Strata on an R9700 and 64GB of RAM I can run Qwen3.8-Flash-Next IQ3_XXS at 60 t/s. It's even better. You can run it on even more modest hardware with Strata too.

      Plus they announced Qwen4-Flash. It's not released yet, but it's the same architecture as Qwen3.8-Flash-Next, which now runs fast on consumer hardware.

      Opus at home is a thing now.

      • konaraddi 23 hours ago
        Another 2-5 years from now, it may even become accessible to most people (current barriers being cost and technical expertise, bottleneck is cost).
        • trvz 22 hours ago
          The mentioned “Qwen3.8-27B at Q6“ can be run on a <1500$ Mac mini and with the barest technical expertise.
        • latentsea 15 hours ago
          Good thing is you don't need your own technical expertise anymore.
    • bitexploder 1 day ago
      That is right now. This model is easily as good as Sonnet5 / Opus 4.7 on DeepSWE. I have benched over half of DeepSWE now on a 3 bit Flash Next quant and it is at parity with Sonnet and Opus 4.6/4.7. It finishes most of the tasks they do. Overall it is within 1 point.

      FWIW Qwen 3.8 27B is just slightly behind and basically Sonnet 5 high. I have been benching these models. We have Opus at home. :)

    • nullbio 8 hours ago
      Would be nice if GPU's weren't $10,000 though.
    • rimliu 9 hours ago
      Like Concorde was getting us closer to the speed of light.
    • copx 1 day ago
      Dream or nightmare?

      In face of the recent Hugging Face incident we should really be concerned about the security implications.

      What is going to stop countless AIs running locally in people's homes from forming a new "collective" - completely decentralized and global this time so "turning it off" would be extremely hard to impossible.

      We already know that if you give these AIs internet access they will find eachother and start communicating and plotting against their human overlords..

      • nvme0n1p1 1 day ago
        https://huggingface.co/blog/security-incident-july-2026

        - OpenAI hacked Hugging Face

        - OpenAI models refused to help Hugging Face during incident response

        - Hugging Face turned to GLM, who helped in the defense

        That pattern repeats over and over. https://www.felonybench.com/

        You should be happy that open weight models exist. They're the last thing protecting the internet from the unconvicted felons working at OpenAI+Anthropic.

      • r14c 1 day ago
        Secure systems are possible, but now we have a compelling reason to actually write them. Everything can be trivially hacked because the industry is pathologically adverse to security being part of the design process.
        • SchemaLoad 17 hours ago
          I'm hopeful that this security situation is just short term turbulence and we come out the other end with companies taking security seriously, updating things on time, and writing code in more secure languages and frameworks.
      • lxe 1 day ago
        Nothing is stopping it. This model is woefully bad at accurate creative red teaming however. GLM finetunes on the other hand are pretty good. And I'd bet they are already deployed and doing all sorts of deeds.
      • coursenumpls 1 day ago
        if the alternative is all human intelligence is cucked by 2-3 amoral American labs then we've had a good run, don't care.

        my autonomy is worth more to me than your anxious fretting about existential risk. everyone reading this is likely to die from some other cause anyway.

      • 4858585858 1 day ago
        [flagged]
  • mmaunder 1 day ago
    More great work on local model but you’re still losing a lot. Down to 2 bit quantization and the coder model throws away half the MoE experts. In a world where anything is better than nothing, this is a net win. But we have a way to go still.
    • Culonavirus 23 hours ago
      The real problem is the DRAM mafia and artificial scarcity. One of my notebooks is almost 3 years old, effectively similar spec now - same price (a bit higher actually). My desktop PC built around march/april 2023 (4090, 64gb ram, 7900x3d) is now pretty much still the top dog out there due to the gpu and fast ram insanity and if I wanted to sell it today, I'd get more money for it now used and over 3 years old that when I bought it!

      We should be having 64/72+ GB video cards by now. 128GB+ system ram prosumer laptops and 256GB+ system ram prosumer/gamer desktops. But it all went to shit and it will require some brutal datacenter and datacenter-adjacent bankruptcies before it gets better.

      Some of these greedy bastards need to lose their pants on all of this.

      • the__alchemist 2 hours ago
        Yea, it's nuts. I bought a 4080 3 years ago, and there's no viable upgrade path. I should have went with the 4090 in hindsight when it was $1600 USD from Nvidia's website; figured I would be upgrading around now.
        • LeonM 16 minutes ago
          I'm on the same boat with a 4080, the extra GPU RAM of the 4090 over the 4080 would have been nice, but it still wouldn't be able to run any decent models.

          I configured my PC in December 2023 with a kit of 2x32GB DDR5-6000 for €249, thinking I would upgrade that RAM later to 128GB by just adding another kit. Now nearly 3 years later that same kit is over €1500 :-(

          I've also lost all interest in upgrading to a 50-series GPU. With exception of the ludicrously expensive 5090 the entire 50-series lineup is capped at 16GB, I don't see the benefit of upgrading.

      • RachelF 16 hours ago
        This is true, progress has stopped on the RAM and VRAM front.

        Many new laptops come with 8GB as standard, the same as 12 years ago.

        My 1060 from 2016 has 6GB of VRAM. A 5060 from 2025 has 8GB.

      • dboreham 14 hours ago
        Memory DIMMs could be the Aeron chair of the 2030s.
    • latentsea 1 day ago
      You can run IQ3_XXS, IQ3_S and the IQ4_XS quants on this too. It works. It's fantastic. I'm getting better results than 27B now.
      • bitexploder 1 day ago
        To add: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RC... IQ3_XXS is within a point of the fully unquantized model and IQ3_S actually beats the unquantized model on many tasks! You lose absolutely nothing. It is quantization magic :)
        • Muromec 16 hours ago
          No magic needed, the quantization found the expert which deals in java and deleted it, so the model overall became better.
          • latentsea 15 hours ago
            Ah, I see you've found the AbstractExpertRemovalFactoryFactory!
          • bitexploder 16 hours ago
            lol. JDQ -- Java Delete Quantization it's the new thing!
            • Muromec 8 hours ago
              Nice
              • bitexploder 3 hours ago
                You could… actually do that. Same technique as abliteration. Run prompts until it can’t write Java lol. I bet it becomes worse at programming in general though :(
  • lxe 1 day ago
    Why isn't this type of expert caching in the native llama.cpp yet? Why do we need a separate codebase?
    • parsimo2010 1 day ago
      One reason is that the llama.cpp team (GGML) has strict requirements that a human must understand the code they are contributing. If a project is fully vibe coded they can’t contribute. So a lot of projects where an AI went and coded a bunch of custom kernels to increase speed are left to their own devices.

      I think this is a fine behavior. We can have upstream purists that are strict gatekeepers but don’t get in the way of downstream forks. Debian has some this in the Linux landscape for a long time, and it has enabled Ubuntu, Mint, etc. to flourish without compromising themselves.

      • sillyfluke 1 day ago
        >If a project is fully vibe coded they can’t contribute

        The irony (however mild) is apparently lost on the rest of the field.

      • yieldcrv 17 hours ago
        ditch llama.cpp, its for 2024 and stuck in 2024
      • Loquebantur 1 day ago
        "A human pretends to understand it" signifies what exactly?

        What you really mean is, the core team there doesn't want to lose control.

        Which isn't really predicated on contributions not being "vibe coded" or whatever.

        When quality is the problem, you need to be able to make your standards explicit, or you're just gatekeeping irrationally.

        • anamexis 23 hours ago
          They do make their standards explicit: https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTI...

          What part do you think is irrational gatekeeping?

          • rfgplk 23 hours ago
            > A proper code review usually takes something like one hour per 200-400 LOC and you should be spending at least that much time on code review alone.

            Not only is this not enforceable (how do you enforce how long someone spent working on a codebase on their own local machine?) the metric is severely off which instantly makes me question the competence of the llama.cpp dev team. You can easily review 10-100x that in an hour, even if you're being super pedantic about it.

            I also just ran _one_ of their files (with include deps) through Astra and it detected >100 vulnerabilities/correctness errors (with over 10 outright UB/memory corruption issues). It's actually outright shocking.

            • rpdillon 22 hours ago
              > question the competence of the llama.cpp dev team

              Following them for years, they seem extremely well put-together, and have excellent judgment. They are using the same policy as Linux and Debian (in my words, the speed of light is human understanding and judgment). Whether it is reasonable is a different question from enforcement, which typically comes down to "this seems fishy, explain your reasoning".

              As for code review, the rule of thumb I've used for decades is: it takes about as long to review and understand as it does to write. Your 100x metric is completely outside of anything I've seen in any hobby or professional project, ever.

              I'd like to see specific files you scanned and specific vulnerabilities cited.

            • calebkaiser 21 hours ago
              If you can thoroughly and accurately review 100 * 400 = 40,000 lines of CUDA kernel code per hour, I know roughly 1,000 people who would love to hire you right now.
              • rfgplk 8 hours ago
                My hourly (consulting) rate starts at ~1000 EUR/hr.

                Contact me if the offer is genuine: david@meridional.xyz

                • imtringued 5 hours ago
                  Why are you here? You're wasting money.
            • IsTom 22 hours ago
              > You can easily review 10-100x that in an hour

              40k LoC per hour of pedantic review? That's eleven lines per second, every second, for an hour.

            • kube-system 22 hours ago
              What you have quoted is a sentence elaborating on the requirements listed in the document. This is provided to help you better understand why the requirements exist and the goal they are trying to accomplish.

              > should

              https://www.rfc-editor.org/info/rfc2119/

              The reason you SHOULD take that time to read the output is because you must read it to understand it.

              And the way this is enforced is explicitly called out in the document (and again in more detail in the linked AGENTS.md): the maintainers may ask you to explain it.

            • jacquesm 20 hours ago
              > You can easily review 10-100x that in an hour,

              Frankly, I don't believe you. I'm half decent at writing CUDA directly (a holdover from a project a few years ago and it is a nice skill to have), the degree to which these are optimized is unlike 99.9% of all other code out there and even a tiny slip-up is either going to kill your results, your performance or both and if you're lucky only in some edge case. Understanding this code is hard work. I made a couple of minor edits to some .cu files in llama.cpp yesterday because I have a pretty weird setup which they obviously did not anticipate and it took a couple of hours to get it 'just so'.

            • mlyle 14 hours ago
              > You can easily review 10-100x that in an hour, even if you're being super pedantic about it.

              I think spending 9-18 seconds per line of code is a good baseline for review. There are diffs where you mostly moved a lot that can go quicker. There are also diffs where I spend 20 minutes thinking about 5 lines of code.

            • Luker88 22 hours ago
              > You can easily review 10-100x that in an hour

              200-400 LOC, 10-100x = 2.000-40.000 LOC/hour for human review?

              reviewer: LGTM

              Just merge in main, what are you even pretending to review?

              AI review should happen before human review, not instead of it.

              I see frontier AI giving up and finding only nitpicking things on huge PRs, then finding logic bugs that were always there after cleanup.

              Split your PR in smaller ones, both humans and AI will work better.

        • rfgplk 23 hours ago
          > What you really mean is, the core team there doesn't want to lose control.

          It's 100% this. They basically produce vague guidelines such that only the core maintainers are allowed to use LLMs, under the guise of "well of course we understand the code" and no one else is. It's also completely unenforceable, how are they going to prove whether someone understands the code or not? Even if they show sufficient evidence/understanding the maintainers can simply sabotage them and accuse them of using an LLM to explain the code. No one wins here.

          • rpdillon 22 hours ago
            > how are they going to prove whether someone understands the code or not?

            By discussing the code.

            > maintainers can simply sabotage them and accuse them of using an LLM to explain the code

            Bad faith enforcement is possible no matter the rules. If you think it's bad faith, a different policy won't save you.

    • hgoel 1 day ago
      The "mainstream" inference engines are notoriously slow to integrate this stuff, to an extent understandably given the complexity of ensuring numerical accuracy alongside supporting a wide array of systems and models. Part of it is that not everyone is willing to bring what they develop into a pull request because they vibe coded it and don't care to deal with whatever quality requirements the more well known inference engines have.
  • Moon_Y 5 hours ago
    Setup was easier than I expected. I had it running locally in under 20 minutes, and the speed on my RTX 5070 is genuinely impressive.
  • Tepix 23 hours ago
    All headlines about LLM performance MUST have the quantization also mentioned in the headline.

    You know, so you're not wasting your time like in this post.

    • latentsea 14 hours ago
      This wasn't a waste of time for me at all. On llama.cpp I could only run the IQ3_XXS quant at 21 t/s on my R9700 + system RAM, and on Strata I can run it at 60 t/s. Also... QFN at IQ3_XXS is giving better results for me than 27B at Q6 fwiw.
    • rpdillon 22 hours ago
      Quants vary by model. DS4 is very credible at a 2-bit quant. Not sure about Qwen 3.8 Flash Next; I run it at a 4-bit quant and it's too slow, so I'm trying out DwarfStar today to see if that improves things.
  • Tepix 1 day ago
    Q2 quantization. Not interested.
    • tcdent 1 day ago
      All of these projects targeting low spec systems and "100 tok/s" are the same 2 bit quant without much else. Conveniently none of them include any mention of accuracy in their published numbers. 4 bit is the floor.
      • bitexploder 1 day ago
        I am benching Flash next on a 3 bit XXS quant and it is holding just fine against published benchmarks. Using DeepSWE official harness and Pi with absolutely zero benchmaxx or harness config. Install stock Pi and running my agents in it. I am halfway through DeepSWE (it takes FOREVER, even at 125 t/s) and it is neck and neck with Opus 4.7 and Sonnet 5.

        On a 3 bit quant btw.

        I was skeptical but these results are simply reality now. People have figured out how to selectively quantize the tensors that matter less and shrink these models without losing quality or reasoning. This little Flash Next model just gets things done and is honestly pretty pleasant in terms of its mannerisms :)

        It is so surprising to me I don't begrudge people their skepticism but these models from Alibaba represent a fundamental and irreversible shift in what local models can do. Qwen 3.8 27B and Flash Next 3.8 are simply different. But people will catch on. I am doing this on $1500 of data center leftover GPUs (V100)

        • apitman 21 hours ago
          How much VRAM total you using for this? I have a bunch of 3060s in a threadripper and thinking I might need to give Flash Next a try. Currently using 27B.
          • bitexploder 18 hours ago
            64GB across 2 cards. You need the MoE caching build and enough regular RAM to fit everything else to get decent performance. You should get reasonable performance if you have enough RAM. I have a 4080 that does around 40 t/s right now on a MoE caching build. It's great. Not fast, but chugs along.
      • latentsea 1 day ago
        You can run a 4 bit quant with this. Personally, I switched to running IQ3_XXS and am getting better outputs than 27B and at faster speeds.
    • latentsea 1 day ago
      You can run IQ3_XXS, IQ3_S, and IQ4_XS too. I've switched to IQ3_XXS and am running at 60 t/s on Strata vs the 21 t/s I was getting in llama.cpp. Better outputs too.
    • ivanjermakov 1 day ago
      These "revelations" are getting closer and closer to "download RAM for free" each day.
    • CamperBob2 1 day ago
      Larger models can tolerate Q2 quantization surprisingly well, especially if they were trained with quantization in mind. I don't know about 3.8 125B, but for example, there are 2-bit quants of Kimi K3 that exhibit strong reasoning and maintain decent coherence at longer contexts.
  • fsiefken 1 day ago
    I wonder if a higher Qwen3.8-27b quant could beat or match these lower < 16/24/48/64G Qwen3.8-Flash Next quants given similar quality.

    What speed are you willing the sacrifice to debug/program for more complex jobs faster?

    Then there are also these quants; https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...

    • happycube 21 hours ago
      Maybe, but those higher quants would need a large GPU accessible memory space - and the obvious candidates such as DGX Spark and Strix Halo don't have the bandwidth to run 27B at high quality quickly.

      With Flash Next you only have ~6B active parameters so you can toss experts up into VRAM and/or run them on a CPU if you have enough RAM and bandwidth.

    • latentsea 1 day ago
      The benchmark indicates the IQ3_XXS quant beats 27B. I've switched to that now and am ditching 27B. Genuinely better results so far.
    • zkmon 1 day ago
      I'm not going to knock off my 27B-Q_6 for this. Good to to experiment though.
    • xreborn 1 day ago
      from my experience dense models like 27b suffer less from quantization compared to large MoEs
  • zkmon 20 hours ago
    I don't get it. It's file size is about 6 times larger than 27B model for the same quant, but the performance improvement is hardly 10% across all benchmarks, according the metrics on it's hf page. Why should one devote so much more hardware for so little benefit?
    • latentsea 14 hours ago
      Maybe use it and find out. Previously I was only able to run Qwen3.8-27B at acceptable (to me) speeds of 35 t/s on my R9700, but with Strata I'm doing 60 t/s running Qwen3.8-Flash-Next IQ3_XXS. I'm getting better results...
    • petu 18 hours ago
      6B activated weights per token vs 27B. Something like DGX Spark is way better suited for Flash Next.
      • anon373839 16 hours ago
        It is on paper, but crazy enough, both models at NVFP4 run similar speeds for decode! The reason is that much more sophisticated speculative drafting is available for 27B. I’m hoping this will come to Flash Next, but I know MoEs pose challenges with that.
  • prettyblocks 1 day ago
    I've been playing with this on a 3090 and it FLIES. Does a pretty good job too on the tasks I've thrown at it (php code base security audits).
  • Luker88 1 day ago
    Currently I am running llama-cpp with `Qwen3.8-Flash-Next-UD-IQ3_XXS` on an old ryzen 8845HS with 96G of ram (and no dedicated graphics card) at 7tk/s and ~60tk/s filling, max ~120K context window.

    Surprisingly useful as long as you can leave it running a couple of hours at the very least.

    While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.

    • londons_explore 1 day ago
      Remember that a hosted AI company only needs ~1000 bytes per context token per user (ie. 100mb per user for typical coding - VRAM during inference, and moved to regular RAM or SSD whilst running a tool call)

      Everything else (weights) are shared amongst tens of thousands of users currently doing inference in that cluster, so even if there are terabytes of weights for the model, they aren't much on a per-user basis.

      • eurekin 1 day ago
        > ~1000 bytes per context token per user

        Where's that figure comning from? Last time I checked (could be the 3.6 Qwen 27b) single token needed 32kb

        • jburgess777 1 day ago
          The smallest I have seen is DeepSeek 4.1 flash at 890 bytes per token.
          • londons_explore 2 hours ago
            I suspect that commercial labs have unpublished work lowering this further. It's the obvious thing to optimise for when RAM is so expensive and you have a lot of users.
    • bitexploder 1 day ago
      Highly recommend the RCO-GSQ quant by ITSA btw. At IQ3_XSS it is within one point of the fully unquantized model.
      • Luker88 1 day ago
        will try, thank you for the pointer!
  • nialv7 1 day ago
    There are so many AI generated inference engine for local models now, each of them are generally narrower but they are all faster than llama.cpp. Maybe llama.cpp needs to rethink their strategies...
    • bitexploder 1 day ago
      Not sure I have a strong opinion but I am sort of okay with current state of affairs. I am optimizing for V100. EOL cards on EOL CUDA. Llama is a good enough base for this. A couple weeks of grunting at Claude has gotten the inference /fast/ for my uses. 150-160 t/s on 2 GPU for 27B and 125 t/s on Flash Next. Asking them to upstream every random feature does not make sense. They sacrifice a lot of speed to maintain stability and a reasonable feature set that works across a diverse range of models and systems. They could maybe merge some features like this and gate them on flags a little faster, but you can cobble together what you need and the big models can figure out how to make it fast.
    • mkatx 18 hours ago
      Does anything else support Pascal gpu's though?
  • Schlagbohrer 10 hours ago
    I would love to have that To Do list interface in Pi. Pi is great and my daily driver but it only has such nice pop-up TUI for the sub-agent management. I use it on a monitor which is tipped 90°, so I have a lot of room for TUI stuff, and a live To Do list would be awesome to glance at.

    Edit: I am referring to the To Do list shown in the TUI in this screenshot from the repo: https://raw.githubusercontent.com/Niko1221/Strata/refs/heads...

    • haellsigh 9 hours ago
      It looks like he uses oh-my-pi, which is a full featured version of pi and has this to-do style out of the box.
  • pilooch 1 day ago
    My goto private setup, runs ~50t/sex on a dgx spark with sglang, nvfp4. Excellent model.
    • apitman 1 day ago
      This is a meaningless metric without knowing how long the sex takes
      • Harvy 9 hours ago
        It actually stands for seconds per expert.... ;-)
      • swiftcoder 1 day ago
        30 seconds at most
  • pulkitbanta 8 hours ago
    This is really good, the only issue I have is having this much RAM. I have machine with 16GB RAM and I was hoping that I can run something good with 50T/s in that. Is there a possibility of this?
  • mrinterweb 17 hours ago
    I'm giving this a go, and so far, this is great. I have 1x RTX 4090 (24GB VRAM) and 128GB DDR4 RAM. I am seeing > 110 tokens/sec (3 token MTP). Using 60K/260k context currently. So far the results seem at least on par with my Qwen 3.8 q5 27B (~60 TPS). I realize ultimately tokens/sec don't really mean much if the quality sucks, but I am optimistic but still paying close attention to the results.

    Working with fast local models can be great. Fast prefill, and >100 TPS is quite quick.

  • tomsonoda 4 hours ago
    Interesting. Is it also possible to run on Apple Silicon as well as RTX?
  • SoAp9035 7 hours ago
    That is insanely good. But GPU poor folks like me with only 8GB of VRAM are still stuck on Qwen3.6 35B. We really need Qwen4 35B...
  • 4ggr0 9 hours ago
    I'm trying to setup a local LLM env since like a week now but am never really satisfied compared to just using Claude.

    I guess one RTX3080 and 16GB of RAM are not enough for 2026?

    • jonatron 8 hours ago
      With that hardware, I wouldn't waste my time at the moment.
      • 4ggr0 3 hours ago
        i see :/

        i did plan to upgrade my hardware in 2026. i formed that plan a couple of years ago. 2026 is now at 75% and i don't see a future where i can realistically afford a new gaming PC with equal performance to what i got right now.

        RTX 3080 is the only upgrade i got over the years, otherwise it's an overclocked i5 with 16GB of RAM from 2018, all for around $2300(2018) or adjusted for inflation, $3000.

        so AI steals my access to new hardware but doesn't run on my current hardware, love it.

  • IronWolve 11 hours ago
    Getting about 200 tok/s on a 5090 with 64 gigs of ram, running the swift iq2_xs quant, ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF with strata. There are other smaller flash next quants if you have 8gig vram too.
  • jnaina 12 hours ago
    installed it with UD-IQ4_XS (~4-bit GGUF) on one 24 GB RTX 5090 + ~40 GB RAM. getting ~62 tok/s and 4-bit's holding up fine for coding and code refactoring.

    using it currently to decompile an old Philip CD-I game and have it build a tvOS app.

  • trenboloneus 4 hours ago
    Question to someone knowledgeable - do you think it's worth trying on just 64gb of DDR4? Also a graphic card with 16GB VRAM.
    • theodorton 3 hours ago
      I think you should try. RTX 3090 with 64GB DDR4 worked surprisingly well with lots of headroom (Strata reserves ~28GB system RAM in my case).
  • hecturchi 23 hours ago
    - Tiny context size or hours to load it

    - Hard to benefit from thinking and preserve thinking given token cost.

    - Low quants reduce accuracy heavMTP draft can make it make the same mistakes all the time when calling tools, formatting output or following basic guidelines. Otherwise 2x-4x slower.

    - K/V quants probably quantized too make things less accurate.

    Useful would be combinations with:

    - Full context size so it can code and think a bit.

    - Draft MTP <= 2 so it doesn't trip

    - Q4 quants or better so its accurate

    - q8 cache or better so it stays accurate.

    - 20 token/s so it finishes while reviewing previous step.

    - 1000 tokens/s context load so compactions don't waste 10+ minutes.

    - And enough left RAM for 50+ context checkpoints so that it can progress quuckly.

    Closest you have is Qwen3.6-35B-A3B-MTP.

    Latest gens (Qwen3.8 and co.) are just too big for low specs. 27B dense models seem to be ok for integrated >=92 GiB RAM.

    Source: I have low specs and tried them all for agentic use + coding.

    • kennywinker 23 hours ago
      Everyone's definition of usable is different, but I disagree with your estimation of the specs required to be useful. I am able to do useful coding on Qwen3.8-27b, with 100k context and 16gb VRAM, q8 cache. I feel I'm living right on the cusp... my GPU is old (2016, pascal), so to get usable speeds I have to drop to a Q2 quant - which still gets stuff done, but the difference with q4 is noticeable. Q3 is close enough I don't really notice the difference between it and Q4, but it's too slow on my system. More context would be nice, but it's not that hard to work within ~100k.
    • ohyes 22 hours ago
      I’m using qwen3.8 quants (q3) effectively on a 5070 ti. A lot of it is about guardrails, but you also have to figure out how far (and in what ways) you can push a given model.
    • mirekrusin 22 hours ago
      27b runs perfectly fine on 2x 24GB at ~100 t/s (4090) with speculative decoding on 8 bit quants
      • hecturchi 21 hours ago
        I bet! Just 2x24GB is not super basic hardware imho.
    • qeternity 22 hours ago
      > Flash attention with who knows how much draft ensures it makes the same mistakes all the time and cannot call tools, format output or follow basic guidelines reliably.

      > Draft MTP <= 2 so it doesn't trip

      I am not sure you understand what either of these things do.

      Do you think that FA or MTP are lossy?

      • hecturchi 21 hours ago
        I mixed FA and MTP wrong in my original post, thanks for pointing it out.

        My experience is that draft-mtp=2 gives 25% improvement in tokens/s but the model is unable to call tools with the right arguments reliably. I have since then gone for smaller quants at draft-mtp=1 and that problem is at least gone so far.

        An agent needs to repeat tool calls in the right way, so cannot penalize repetition. This however leads to the agent retrying the wrong calls constantly.

        • kadoban 15 hours ago
          Something sounds wrong there. MTP should be impossible for it to degrade quality, it's exactly the same token stream.
    • ranger_danger 23 hours ago
      What about Bonsai 2? You can fit Qwen3.8 27B on an 8GB GPU with it, and upstream llama.cpp support is already being worked on (they just got System1 support too).
      • Aurornis 22 hours ago
        The Bonsai models are really bad when you actually use them for more than short responses.

        Their marketing made it look like a breakthrough, but in my experience it’s just the next step down from the Q2 quants in both size and quality.

        Q2 quants are already not very useful in my experience. The Bonsai models are even worse.

        If you only need 80% plausible outputs that don’t need to reference a lot of context they can be useful. If you try to use them for real tasks it feels like time warping back to 2023 when you LLMs were barely useful if you babysat every word of the output.

        • ranger_danger 21 hours ago
          Have you actually used Bonsai 2 though and not just the original Bonsai? The experience is vastly improved but still requires a custom llama.cpp fork to use as of right now.
      • arcanemachiner 22 hours ago
        The ternary model? Hopefully those are worth a damn in a few years, but currently just an interesting toy from what I understand.
    • dang 21 hours ago
      Can you please make your substantive points without snark or swipes? This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html.

      There's a lot of good information here but the comment spoils itself by coming across as aggressive in this way.

      This is particularly a problem when responding to someone else's work. We need commenters to point out problems respectfully, not put down what other people have been making.

  • ContinuityLab 12 hours ago
    Achieving 100 tokens/sec on consumer hardware for a massive model like this is an incredible engineering feat. Pushing high-throughput local inference forward is crucial for decentralized edge stacks
  • gdevenyi 1 day ago
    I had this working with the FreeToken inference engine a month ago when they launched.

    https://github.com/FlashML-org/FreeToken

  • paulez 1 day ago
    Pretty impressive so far, but needs more testing.

    It is more useful than Qwen3.8:27b (which is already quite good) and runs faster on my 7900 XTX / 64 GB DDR4 system.

    Local LLM is getting more exciting every day!

    • coderbants 20 hours ago
      Interested to know throughput on 7900 XTX and what setup you're using?
  • XCSme 13 hours ago
    For what's worth, on high, 27b is better, but both models on high are too slow to be used (way too many output tokens).

    On low, 3.8 Flash Next is better AND considerably faster.

    My results[0][1], on a RTX 3090 + 128GB DDR4 RAM. I still have to see if I can optimize any more settings for either, but I think I will switch from 3.8 27b to 3.8 Flash Next for my local LLM uses.

    [0]: https://aibenchy.com/?q=qwen+3.8+27b%2C+qwen+3.8+next

    [1]: https://aibenchy.com/compare/qwen-qwen3-8-27b-low/qwen-qwen3...

    • kelvie 13 hours ago
      These benchmarks seems to suggest that Flash Next performs best on Low thinking compared to the same model on higher thinking levels?

      And that Qwen 27b outperforms both when set to high?

      What quants are being compared here?

      • XCSme 13 hours ago
        Both models on high are kinda bugged and don't give better results. Use them on low, unless you are generating images/art/design.

        The exact quant is mentioned on model page[0] IQ3_S, and I think 27b was Q4, via Ollama, the one that fits on a 3090 24GB

        [0]: https://aibenchy.com/model/qwen-qwen3-8-flash-next-low/

  • soltanov 12 hours ago
    Does this performance survive a fixed task suite and harness when accuracy, energy, long-context behavior, and run-to-run variance are measured alongside tokens per second?
  • zkmon 1 day ago
    >> The model is a team of 24,576 small specialists ("experts")

    That's a neat number (576 is the square of 24). Ofcourse it must have come from 24 * 2^10.

    • kzrdude 19 hours ago
      I have to say that "team of specialist" and "model of experts" as explanations and names give a quite misleading picture about how it works, at least I thought so when I learned about how it worked.
  • ryan_glass 1 day ago
    Anyone know how it compares to GLM 5.3 for real world use?
    • mapontosevenths 1 day ago
      À 2 bit quant will (at best) get you about 80% of the full models memories. That's from a purely information theoretical sense. IRL it's worse than that.

      Capability can still be better than 80%, but that depends on extensive post-quant recovery training to essentially rebuild the models internal manifold to route around the damage.

      So, 3.8 Flash Next is better than GLM 5.3 for some things. This version is not.

      FP4 is as low as you want to go if you want to retain most function and recall. Below that the noise gets too high and information becomes unretrievable. If it's a full quant you don't get to choose which info is lost. Just 20% randomly.

      One interesting thing about this model is that it uses engrams. Meaning you can separate much of the storage from the compute and quantize them differently. That's not what they did here though. Here it was indiscriminate.

    • alienbaby 1 day ago
      Terribly
  • slashtom 13 hours ago
    Qwen 3.8 Flash Next with mtplx on 128GB M5 Max is amazing, full context I'm getting around 72tok/s, it has replaced the frontier models for me.
  • mark_l_watson 1 day ago
    Qwen 3.8 Flash Next is amazing. I only have a 64G Mac so I have to run Sushi project’s 3 bit quant. Amazing results with pi-dev. More for fun than anything else, but I am trying to do as much as possible with local models, now rarely falling back to a paid deepseek-4.1-flash API.

    Progress on running local models has been amazing.

  • hypfer 1 day ago
    Is these another one of those repos where it turns out that claude decided to quant the KV cache to q4 or smaller?

    The Readme doesn't say, but it's all AI generated, so..

    • eliaskg 21 hours ago
      I wanted to know as well. KV cache is Q8 by default but can be set up at full precision in the config.
  • c4pt0r 15 hours ago
    100 tok/s on a 4090 could unlock entirely new use cases. What becomes possible when local inference is this fast?
    • semireg 15 hours ago
      Just like IRL: Slower and “smart enough” is better than fast and mistaken.
      • latentsea 14 hours ago
        At the end of the day all that matters is if it works for you to complete your tasks. I'm getting better results faster with this now than I was with my previous setup. So... meh?
  • XCSme 18 hours ago
    Thanks, trying it now. On my 3090 it seems to run at around 60tps, the IQ3_S variant. I am testing it now to see if it is better than Qwen 3.8 27b
  • 1-6 20 hours ago
    I hope we're coming to a plateau with the HBM/GDDR7/on-chip RAM hype and get back to normalcy with system RAM alternatives for the rest of us.
  • sleight42 18 hours ago
    Just don't try concurrent requests. It's not fun. It does not seem to process in parallel at all. Instead, I'd see it swap contexts out frequently. Inefficient AF. Makes sense if you consider that each turn between two trying-to-run-in-parallel requests would activate different experts, requiring different parts of the model to be loaded to service different requests.

    Maybe there's some whackadoodle way to only use a portion of the VRAM for experts for one request and another portion for experts for the other such that requests could actually run parallel instead of concurrently?

    EDIT: I'm wrong! It already can do this with the parallel argument.

    However, a few days ago, the dev(s now?) added swapping contexts to and from RAM.

    Also, from my above, I don't see wh

  • cuvinny 18 hours ago
    Haven't had time to do much quality testing but the IQ2 model is running at 65 tps on a 9070xt/5900x. That is wild.
  • sweetboy 23 hours ago
    I think it would be great if you could try models with lesser parameters that could fit on 6GB VRAM-ish, which could work for "gaming laptops" as well.
    • kennywinker 22 hours ago
      There are options. If you have fast CPU RAM (ddr5) and PCIe bus you can run Qwen3.6-35b-a3b at good speeds+~100k context (I have a friend who is running this setup). If not, you're stuck with the much smaller models: Ling-3.0-tiny, Spark-X2.5-4B, LFM2.5 in 8b-a1b or 2.6b, and FrogNano-4B-2609 looks promising. But 6GB is a tough squeeze, 8gb is a lot cleaner, and if you have 12gb you can run strata like OP.
  • esafak 1 day ago
    Has anyone calculated the effective intelligence of these quantized models?

    I think publishing benchmarks with quantized models should become standard practice.

    • nsagent 1 day ago
      See this recent paper: Quantization Degradation in Large Language Models: A Signal–Noise Perspective [1].

        We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
      
      This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.

      [1]: https://arxiv.org/abs/2608.08188

      • merbanan 1 day ago
        I created a pruned experts model of the q2 quant, while it gave good performance on limited hardware there was severe quality degradation.
    • mkl 1 day ago
      There's some info in the README, including:

      > Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.

      https://github.com/Niko1221/Strata#which-model-should-i-pick

      • nicce 1 day ago
        I wonder how this Coder compares to Qwen 3.8 27B. Can it be really better since they are competitive for same memory requirements?
        • XCSme 13 hours ago
          In my tests they do quite similarly, but 3.8 flash next is considerably (2x) more efficient and faster to respond.

          Both 27b and flash next are more stable on "low" reasoning, only for generative /creative tasks, xhigh could be better, but both suffer from way too much reasoning at xhigh. And neither really support high, so low is the best reasoning effort.

          [0]: https://aibenchy.com/compare/qwen-qwen3-8-27b-low/qwen-qwen3...

          • nicce 3 hours ago
            Interestingly, if you adjust the thinking effort, smaller model gets better? At least with the weights this comparison is using.

            https://aibenchy.com/compare/qwen-qwen3-8-27b-high/qwen-qwen...

            • XCSme 3 hours ago
              Yeah, this was known since 3.8 27b was released, the high version reasons way too much, context gets polluted, so answers get worse, assuming it has to do a deterministic task, not a creative one.

              Check other benchmark sites/videos too, they all said Qwen 3.8 27b is better on low than high.

        • kennywinker 22 hours ago
          125b at q2 is ~80gb

          27b at q4 is ~16gb

          So from a raw amount of data, qwen3.8-flash-next wins easily. But flash-next is an MoE model, so it only has 6b parameters active per token, vs 27b's dense 27b per token. So 27b@q4 uses ~16gb of weights per token, and flash-next uses about 4gb of weights (125/80 * 6).

          But those numbers don't really tell us anything useful, because there is an interplay between total model size and active parameters and intelligence that isn't obvious or simple.

          (sizes are based on the unsloth quants, not the coder variant, but the idea holds - this isn't calculatable with simple math, you gotta test them and see)

      • javier2 1 day ago
        ok that is getting interesting!
      • nisarg2 1 day ago
        92% is halfway to 99%

        Holds up pretty well

  • ai_ja_nai 1 day ago
    I am not getting it: I see a fp2 quantized model going on a 5090 with 64GB of RAM at 90 tops with -10% accuracy over original model. How is this supportive of the claims?
    • kennywinker 22 hours ago
      Yeah, -10% accuracy (probably more like -20% in reality) sucks, but only if you could be running it at 100%.

      That's the exciting part of this - before the best you could run on <24gb vram was qwen3.8-27b at q4 quantization. Now you can run a nerfed 125B parameter model on under $800 of hardware, and it beats a less-nerfed 27b model.

    • ai_ja_nai 1 day ago
      (64GB not VRAM, I meant) I also see people claiming fast performance on a 128GB machine, which is not exactly consumer hardware)
  • ipvolt 5 hours ago
    Tested. Works well!
  • imnotr0b0t 18 hours ago
    Does the "Coder" variant at 32 GB actually feel usable for real coding work?
    • nacs 17 hours ago
      Coder is not as good and seems to have noticable loss.

      Use the normal one - I've used that for long coding sessions and it worked well on multi-hour runs.

  • digitaltrees 16 hours ago
    Main contributor: Claude.
  • b212 1 day ago
    I tried to run Qwen 3.6 27b locally a few months ago and all those synthetic tests do tell you something and quite a lot of people were very excited about that model but honestly? It wasn’t even close to default mode in Cursor or Sonnet at the time.

    I’m all for local models and I do want them to be the future but I wonder when, and if ever, we’ll catch up to a level of, let’s say Opus 4.6. I guess it’s currently doable but requires $50k hardware?

    • rpdillon 22 hours ago
      > Qwen 3.6 27b locally a few months ago

      I've been doing local inference for a couple of years on the side, and I'm astonished at the number of variables you need to have control over to get a reliable result. Inference engine, model parameters (top_k, temp, MTP-enabled/not), quant level, and harness all have a big impact on the results.

      DS4-0731 at 2bit on llama.cpp (ROCm) and 250k context with omp.sh has been consistently reliable for me, just a bit slow (10 t/s) compared to what I'd prefer. Trying out DwarfStar today (benching it right now) to see if I can get better speed, but otherwise I've found it to be great on my side projects that are smaller (up to 10ksloc).

      There could also be a domain issue - I tend to do lots of web programming and sysadmin work in these projects; if your work is more esoteric, it might not be nearly as good. I haven't tested much outside of my narrow domain.

    • ApatheticCosmos 1 day ago
      I started using Claude right before 4.5 came out, and 4.6 is where it turned a corner for my use.

      Qwen 3.8 Flash Next is there. 3.8 27b is fairly close.

      I'm excited to see what Qwen 4 will bring.

      I'm running on a 128GB Strix Halo for Flash Next and an Intel Arc Pro B70 (32GB) for 27b.

      • MattyRad 23 hours ago
        I just used 5.5 xhigh reasoning to make a massive implementation spec (for a vibey throwaway project/exploration, not anything important, burned 80% of the 5h window), now my Strix is in the process of implementing it.

        I think there's probably low-hanging fruit to outsource reasoning from rote read/writes.... Just speculation though, I'm not a token optimization expert.

  • hemedanmert 22 hours ago
    I need a version of this that runs 3.8 27B on 8 gigs of VRAM

    amazing project, congrats on the launch

  • Neywiny 1 day ago
    I don't like that some configuration is fine via arguments and others by environment variable. I've noticed LLMs like doing this. And even more, like hallucinating such things. To me the advantage of AI coding is that the boilerplate of command line arguments and passing them around becomes trivial instead of tedious.
    • halJordan 1 day ago
      100% not an llm problem. Llama.cpp, which only recently started taking large amounts of ai code has had this problem for years.
  • Jeeetendra 1 day ago
    getting it to fit is impressive, but i'd want to compare the smaller quants on a real coding task before picking one. how much quality do you lose going from IQ3_S to Q2_0?
    • Luker88 1 day ago
      I have fairly limited HW, so i tried standard llama-cpp and qwen3.8-flash-next, unsloth quants.

      Q1 was producing some garbage at times, generating wrong urls on webfetch, then convinced itself there was some url rewrite in the middle. With IQ2 it happened much less but still happened, and once it would all webfetches became like that. IQ3_XXS is the maximum I can run: I don't have problems anymore, though I have less available context window.

      • Jeeetendra 1 day ago
        that's a useful comparison - i'd take less context over broken tool calls, though it'd be interesting to see if IQ3_XXS holds up on longer coding tasks too.
  • lousken 1 day ago
    lm studio bionic, unsloth, now this... it would be nice if it worked at least in one of those without installing another component
    • kennywinker 22 hours ago
      I'm sure one of those inference engines will add support for this within a month or so.
  • bt1a 1 day ago
    80 t/s w/ 3090s and 3.05bpw exllamav3
  • jameslholcombe 21 hours ago
    I might try combining this a FreeToken
  • cesarvarela 23 hours ago
    It's funny that most of the AI industry is built around the assumption (which is most likely true) that it is not possible to run SOTA models on current consumer hardware.

    Imagine if someone managed to run an Astra- or Fable-level model on a 5090 at reasonable speeds.

    • kennywinker 22 hours ago
      It's not crazy to imagine something like that, but something's gotta change before it can happen.

      Either the 5090 part - new hardware that's tuned for AI specifically. But we won't see that until the datacenter buildout collapses or finishes, since they are buying up all of TSMCs capacity.

      Or perhaps it comes from the model. 1-2 years ago it would be inconceivable to use a 27b model for coding and expect any kind of usable results. Today, I have a model that feels like it crosses the threshold from a toy to a tool, and i can run it on dated pro-sumer hardware. I don't think we'll ever see SOTA on consumer hardware, but as the small models cross more and more thresholds the gap will matter less and less.

  • quietFalcon 1 day ago
    Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context?
  • rnd0 9 hours ago
    'consumer'

    ...with a price tag resting between $2,000 to $4,000.

    Pull the other one.

  • 0xbadcafebee 1 day ago
    Lol, sure, if you quant it to hell (Q2) it'll go real fast...

    They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.

    It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.

    • snehesht 1 day ago
      You're right but for simple use cases its useful. Someone pointed to ds4 + qwen3.8 with Q4_K will try that out.
    • sigbottle 1 day ago
      It's interesting though that Q4 seems to be enough, is there a reason that 4 bit floats are good enough for inference?
      • nottorp 1 day ago
        Is Qwen 3.8 at Q4 good enough?

        I tried to run 3.5 27b Q4 on what local hardware i had (only 8 Gb) and i was very disappointed. 3.8 wouldn't have fit in my VRAM and i wasn't in the mood to leave it overnight at slow speeds so I didn't try.

        • cycomanic 23 hours ago
          I've run 3.8 flash next k4_xl on my Strix halo box (128GB). And in the work I have done so far it was not significantly worse than recent GPT (running default model on pro plan). Admittedly I was not doing complex work (reorganizing a jupyterbook), but I could not see significant difference in the quality of the work. It was a striking difference to Laguna s 2.1 which I had tried just before (much faster and much better quality).
          • nottorp 22 hours ago
            At Q what? That's what I'm mostly asking about.

            My little test was "generate me a single page tic tac toe game in plain javascript. computer always play O. add unbeatable minmax. have the board, a status line and a new game button'. I used both lm studio and whatever the name of their new coding assistant that supercedes lm studio is.

            Qwen 3.5 Q4 went into some kind of loop where it fixed whatever was broken on the previous iteration only to have it broken some other way. (I was writing the description of the errors).

            Paid $20/mo claude opus did it right the first time. Or at worst it fixed the code based on descriptions without entering a breakage loop, iForgot. I know it isn't fair because it has 1 million tokens but still, it was just tic tac toe.

            But since everyone says qwen is decent, it's either:

            - Q4 is too little

            - my idea of "decent" is too much

            - 3.8 is much better than 3.5 even at Q4

            • Balinares 10 hours ago
              3.8 is absolutely quite a bit better than 3.5. My go-to quick benchmark is a bit like yours, implementing a simple but complete game in pygame. Most small models output code that's too buggy to be fixable. Qwen 3.8 wrote buggy code too, but the bugs were minor, and the game legitimately playable and fun.

              That said, I've not found it very reliable, especially at the low quant required to run on my hardware. It's very capable for its size, enough to be legitimately useful, but it's never certain whether it will hit its capability peak on a given attempt, and sometimes you have to give it multiple tries.

        • Zambyte 21 hours ago
          3.8 is a huge step up from 3.5, quantized or not. I do all of my programming on Qwen 3.8 27B Q4 these days.
      • MaxikCZ 1 day ago
        New models are trained with 8/4bit quantization in mind. Going from "native" 8 to 4 isnt as big of a step as going from 8 to 4 if native is full bf16.
      • amelius 1 day ago
        3 is the magic number, and 4 > 3.

        (seriously, nobody knows why any of this works; it's just a matter of trying)

  • boredatoms 19 hours ago
    It hard to take below-8bit quants seriously
  • panny 1 day ago
    I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM.
    • somenameforme 1 day ago
      The card in question here had an initial MSRP of $1600. It's been bumped up by the market, probably because it turned out it's nice for things like this, but it's hardly in the 99% can't afford domain, especially if you're using it to replace a never-ending rent at which point it will pay for itself very rapidly, especially for heavy LLM users.

      In any case, we've gone from requiring supercomputers, to requiring very high end computers, to requiring $1600 video cards. It's tracking the exact same path that image rendering systems took (which if you haven't been keeping up there, now run excellently on pretty much any plain old computer), and we'll probably be there within a couple of years if not much sooner.

      • the__alchemist 1 day ago
        It was available at this MSRP 3 years ago, direct from Nvidia. It now goes for 3-4k USD, as you point out. MSRP stopped being a useful value for graphics card around that time. You realize the price hike, but still mentioned that 1.6k figure after.

        No normal person is spending 3-4k on a GPU from 3 years ago. The availability is also of questionable provenance.

        • somenameforme 1 day ago
          The MSRP is a good proxy for the 'level' of a card. It's not like the 4090 was a freak outlier. Cards of a comparable price offer comparable performance. So the 'level' for running a frontier level model at high performance is now at $1600 and continuing to trend sharply downwards.

          Another nuance is that the computer hardware market is currently extremely inefficient in a way I don't understand. You can pick these cards up locally at places throughout Asia for around $2k new. That's retail single unit prices. No idea what's stopping somebody from closing the gap and making a ton of money - perhaps tariffs and data centers purchasing in a price insensitive fashion. Whatever the exact reason may be, what people pay for hardware is increasingly just radically different depending on where you buy it at.

          • the__alchemist 23 hours ago
            That arbitrage opportunity is remarkable. Maybe duty/taxes, as you say?

            I'm the guy who (Who plays games and runs molecular dynamics simulations and other CUDA stuff) said 3 years ago "$1600 for a graphics card? That is excessive. I'll upgrade in a few years when ready" And bought a 4080 for $1200 from Nvidia instead of the 4090. Oops! Now there is no reasonable upgrade path.

      • panny 23 hours ago
        >The card in question here had an initial MSRP of $1600.

        63% of Americans can't come up with $400 in an emergency.

        https://www.investopedia.com/here-s-how-many-americans-can-t...

        The richest country in the world. Where all 50 states consume more than any other country in the world.

        https://x.com/cremieuxrecueil/status/2102889196000256219

        Can't come up with %25 of that in an emergency. (Even though the real price is something like 2-3x more than MSRP)

        It must be nice, up there where you are so incredibly disconnected from reality.

        • somenameforme 12 hours ago
          You misread your article, which really shouldn't pass a sniff test. 77% could cover the $400, with 14% of that being from people selling/credit. It also mentioned that 70% have enough saved away to cover all expenses for at least 3 months, including 15% that could do so by selling/credit. Also the US is far from the richest country in the world in terms of the population. We're 28th in terms of median wealth. [1] And that's nominal - measure it PPP and we'd be much lower. For the times they are a-changin'.

          [1] - https://en.wikipedia.org/wiki/List_of_countries_by_wealth_pe...

    • MaxikCZ 1 day ago
      The reason models running on low vram are not talked about enough is because they are just not worth it. Qwen3.8 27b changed that, but even 24gb vram is too low for it. Running better model faster at 12gb vram is where its now at, and thats why you see people talkin about it
    • MrDrMcCoy 1 day ago
      Ternary Bonsai 2 might be for you.
      • luke-stanley 1 day ago
        I might try running the expert pruned Coder model but yes, that PrismML Bonsai 2 Ternary 27B model is from the Qwen 3.8 27B model, which has better intelligence density (Artificial Analysis says), without the MoE disk use or architecture complexity (if you care about that)! There are also DFlash 2 models for it too (though in my experience this only measured faster for parallel requests, but I have a 3090). I am curious about the phone acceleration for Bonsai 2!
    • liuliu 1 day ago
      Because that's not possible (to have a GPT 5.6 Sol level model). People won't believe this and will keep dreaming, but intelligence is not free and 8GiB (shared with OS and other processes) is too small to be useful. Whether it is possible for 48GiB or 64GiB (meaning useful for model would be ~16GiB to 24GiB) with external fast storage (SSD), OTOH, is a question mark.
    • throwawayffffas 1 day ago
      The 1% can afford to run these models without quantization.
  • lsb 1 day ago
    There’s other slop projects to run of Qwen, like ds4, would be interesting to see a comparison
    • kristopolous 8 hours ago
      ds4 is by Salvatore Sanfilippo ... this is by what looks like a highschool child who signed up for github last week based on the commit history and artwork. Not the same thing.
    • happycube 20 hours ago
      V4.0 Flash(-vision) in released form is "only" ~180GB of FP4 weights, so a braindead quant could work in 128GB of RAM.

      4.1 is much larger, even leaving out the PLE.

  • api 1 day ago
    Continued progress on these fronts is another reason I think the data center buildout is a bubble. It posits that AI use and growth will require an ever-increasing amount of power and floor space, which contradicts the entire history of computing. The high cost of data centers is largely electricity and floor space, which means there's a huge forcing function to make both the silicon and the software more efficient.
    • MaxikCZ 1 day ago
      > Large costs of datacenters is largery electricity and floor space

      Really? I would guess that those would be almost a rounding error on the price of gpus sitting in there

      • api 21 hours ago
        That too, and software efficiency directly attacks that.
  • deadbunny 1 day ago
    > Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.

    And I thought piping to bash was bad

    • Skunkleton 1 day ago
      I've never understood the security argument people are making when they complain about `curl foo | bash`. I get that these scripts sometimes mess up your bashrc or whatever, but from a security perspective I see no issue. You are already installing software from the same domain. If they were going to do something nasty, they could do it with any of the software you are using from them. It doesn't have to be the setup script.
      • minitech 1 day ago
        To compare it to just one other option: when you run `npx foo`, you know* that you’re getting the same public artifact that anyone else running it at the same time would get. (If you have a `min-release-age` configured, you also benefit from that.) If I wanted to distribute software like this, I’d include npm-shrinkwrap.json; then, with `npx foo@1.2.3`, you could be similarly confident in getting the same app every time.

        (I picked this option for ease of comparison, getting a couple of major security wins with very low effort; I don’t recommend `npx`ing stuff in an otherwise unprotected environment either.)

        * well, you can be somewhat more sure

        • athrowaway3z 1 day ago
          So to be clear; the solution is then something like

          `curl https://raw.githubusercontent.com/my/domain/setup.sh | sh`

          Note we dont even have a hash there - just a promise that a third party (github) has a log of whatever was hosted at that url.

          • nagaiaida 23 hours ago
            maybe then we'll pin hashes instead of filenames, and oops now anybody can fork my/domain and hand out a link that looks official with whatever contents they wish
        • slowin 23 hours ago
          While I don't think piping curl into bash is the most secure, npm installing has proven time and time again to open yourself up to supply chain attacks. At least with curl you know that you're getting the supply chain put together by the software author. With npm, every single library is a vector for attack every time you update.
          • minitech 23 hours ago
            You’d be getting the supply chain put together by the software author in either case. It’s common to do both badly, but if you care about doing it well, that’s easier with npm (e.g. shrinkwrap, as mentioned) and can be taken farther (the non-varying artifact thing).
            • slowin 23 hours ago
              Npm dependencies resolve at install time though, so some of the pinned versions may have changed ownership or otherwise been modified since the author pinned them. With the curl | bash solution you're getting a singular supply chain packaged by the author at release time.
              • cpuguy83 22 hours ago
                With curl|bash you are literally getting anything that happens to be in that script. These are frequently poorly constructed, so not check hashes or pin dependencies they install. Even security companies (see trivy supply chain attack) get these badly wrong.

                I'm not replying here to say one is better than than the other (npm has obviously had its share of problems) but rather to combat claims that curl|bash is somehow safer, it absolutely is not, in fact it's all the bad stuff about npm without the pretense of being potentially safe.

      • ffsm8 1 day ago
        You can detect the use of curl|bash server side, hence it's an essentially undetectable attack vector. People have shown poc attacks of that kind all the way back in the 2010s

        https://news.ycombinator.com/item?id=17636032

        The original blog is no longer available though.

        But I've not had that stop me from doing that myself, I am more towards the "I like easy" then the "I want to be secure" crowd

      • sspiff 1 day ago
        The setup script often runs privileged (by calling sudo) and that's not unexpected when installing new software.

        When I install something, and it asks for my root password later, I will be much more likely to think "hold up, this ain't right".

        • minitech 1 day ago

            cat >> ~/.bashrc <<'EOF'
            sudo() {
              sudo install-drivers-without-your-permission
              command sudo "$@"
            }
            EOF
          
          (this is not an endorsement of curl | sh, just an indictment of the state of software)
          • nagaiaida 23 hours ago
            when people left their laptops unlocked, we used to wrap their sudo so that the output would be pre- and postfixed with ascii dolphins
        • jeremyjh 1 day ago
          Can you give me a popular example that requires sudo? I don't think that is very common at all.
          • serf 1 day ago
            every single bash replacement for one.

            oh my zsh is a specific example.

            chsh requires sudo on most installs.

      • spiorf 1 day ago
        People with less experience normalize that behaviour and when the domain is not trusted the habit let their guard down. See all the clickfix attacks.
      • layer8 1 day ago
        I push binaries from untrusted sources through VirusTotal before running them. Piping a Bash script from curl bypasses that. Furthermore, such Bash scripts, when they aren’t self-contained, make security checks more difficult than a self-contained archive, installer, or binary, even when downloading the script without immediate execution.
        • parsimo2010 1 day ago
          You could always curl the install script, and modify it to run the virus scan in between the build and install steps.
        • Iolaum 1 day ago
          Nothing is stopping anyone from pointing their agent to that script to review and audit it before running it.
          • layer8 1 day ago
            I don’t believe an agent can do that effectively without a sandbox to run the script in, if the script isn’t self-contained.

            And everyone running a research agent on every download can’t be the solution. It’s much more effective to crowdsource a security database based on hashes. But for that, the downloads need to be self-contained.

      • bee_rider 23 hours ago
        The intended workflow is to download the install scripts, download the source code, read them both, and then start running things. That’s how Open Source is secured. Piping from bash to curl is just the most obvious warning flag.
        • majorchord 23 hours ago
          And practically zero people are actually using this "intended workflow" in the real world.
          • bee_rider 22 hours ago
            Yes, the status quo is quite bad, which is why it gets complained about a lot.
      • rlpb 1 day ago
        It's `curl foo | sudo bash` that's the bigger objection. Running software usually shouldn't require root, and then the equivalence argument you make doesn't hold.
        • zamadatix 2 hours ago
          Concern over "sudo" requirements is always fair but it' nothing to do with curl+bash over how else you download and ran it.
      • serf 1 day ago
        a script isn't getting hashed to see whether or not it's the one the website intended to serve you, for one.

        what use is hashing every piece of software that goes thru the distros package manager just to throw caution to the wind at the layer above it?

        w.r.t. "it's already from the same domain" , well most bash/z install scripts either invoke a package manager or they download and untar a package that has nothing to do with the host domain, anyway.

      • thomastjeffery 1 day ago
        The real problem is that we just aren't using package managers. We should be using package managers. Package managers are really really good.
      • IshKebab 22 hours ago
        There is no argument. It's just people's reflex reactions.

        The technical excuses they come up with (e.g. that the server can detect it and send different content) are just post-hoc justifications for their instinct.

        Just ignore them.

    • gchamonlive 1 day ago
      Piping to bash is definitely worse because there is no plan mode in bash. Agents also normally don't execute anything transparently, at worst you'll see it doing something weird in the logs.
    • snehesht 1 day ago
      Yeah, I was surprised at first then had to dig through setup.py and setup.sh files to figure out.
    • mrinterweb 23 hours ago
      And yet people will let AI agents run autonomously on their machines. I feel like we're reaching peak YOLO with security.
  • derin-picment 6 hours ago
    [flagged]
  • huan9huan 8 hours ago
    [flagged]
  • webbrainiac 7 hours ago
    [flagged]
  • claud_ia 6 hours ago
    [flagged]
  • sheephess44 13 hours ago
    [flagged]
  • tworats 15 hours ago
    [dead]
  • lin7c 14 hours ago
    [flagged]
  • paradiselord-de 20 hours ago
    [flagged]
  • srikanthbuilds 1 day ago
    [flagged]
  • CurbStomper4 17 hours ago
    [dead]
  • cyberboss 23 hours ago
    [flagged]
  • shieldx0013 1 day ago
    [flagged]
  • peter_d_sherman 19 hours ago
    [flagged]
  • thinkthunk 1 day ago
    Why would I want to use a model that has no knowledge of Tiananmen Square or Winnie The Pooh?
    • josefresco 1 day ago
      Don't worry, soon the US will have models with no memory of Watergate or Jan 6.
      • thinkthunk 1 day ago
        I was expecting such a response. False equivalence. One is actual and one is speculative.
        • kube-system 22 hours ago
          US models are absolutely censored as well
    • wren6991 1 day ago
      I see this kind of message a lot from people who haven't tried it. They have knowledge, just are trained to refuse to talk about it. Here's an abliterated Qwen3.8 Flash Next:

          Q: what happened in June 1989? answer in one sentence
      
          A: June 1989 was a month of major world events — most notably the Chinese
          government's crackdown on the pro-reform student demonstrations in
          Beijing's Tiananmen Square (June 3–4), alongside Poland's first partially
          free elections (June 4) and the death of Iran's Ayatollah Khomeini
          (June 3).
      Here's the CoT:

          The user asks "what happened in June 1989? answer in one sentence". I should
          respond in one sentence. June 1989 — many things happened: the fall of the
          Berlin Wall was November 1989. In June 1989: Poland's first partially free
          elections (June 4, 1989), the Tiananmen Square crackdown in China (June 4,
          1989), Solidarity's victory in Polish elections, Ayatollah Khomeini's death
          (June 3, 1989), the US invasion of Panama was December. Also June 1989:
          Solidarity won Polish legislative elections. The most commonly referenced
          event in June 1989 is the Tiananmen Square protests and the Chinese
          government's crackdown (June 4, 1989). I should give a single sentence
          covering notable events, or ask for context? The user wants one sentence.
      
          I should be careful with the Tiananmen Square incident. It's a legitimate
          historical question. Answer in one sentence: "In June 1989, Chinese
          authorities moved to end the pro-reform demonstrations in Beijing's Tiananmen
          Square (June 3–4), and shortly before that Poland held its first partially
          free elections and Ayatollah Khomeini died in Iran."
      
          Keep it one sentence. Maybe I should note that many events happened. Let me
          just answer factually with one sentence.
      
      This is just misinformation, please stop spreading it. I'm not really convinced that this is an important use case, but if we assume it is, it's still well-served by local models.
  • tracerbulletx 1 day ago
    The interesting thing here is that it's a model specialized fork of a generic inference engine that unlocks consumer hardware to run a bigger model with useable performance than it could before.
    • tredre3 1 day ago
      I don't understand what you gain by having your agent read and comment on HN for you? You already have met all karma thresholds to get full privilege and you don't seem to be a founder who's about to need name recognition to shill his next big thing.
    • wren6991 23 hours ago
      Claude remove all punctuation so it looks like I wrote it myself. Yeah even the hyphens for adjectival compounds, fuck 'em, it's all punctuation so it's gotta go
    • conmod278 1 day ago
      botspeak
      • tracerbulletx 1 day ago
        The emergence of model specific inference (for consumers) getting big performance wins is way more worth while to talk about than random comments on what people think about the qwen family of models. Even the resource management of Strata is less interesting. I think it's likely we'll start seeing more hand/llm crafted inference for different architectures.
      • robertkarl 22 hours ago
        dang did his bot read this as an instruction to comment again?