Nvidia DGX Spark as a daily driver

(daniel.lawrence.lu)

70 points | by plun9 3 days ago

17 comments

  • InTheArena 10 minutes ago
    I have a DGX and a Ryzen AI Max 395 - while I love both of them, there are a few critical things that leave the DGX in use, while the Ryzen "just" is my primary homelab server. The biggest thing is prefil numbers, and the performance impact of higher context sizes. Qwen 27b is a great model, nemotron is decent, gemma is workable. But all of them need reasonable context for reasonable outputs.

    Unfortunitly, as others have noted, the DGX OS experience... sucks. My hope is that the RTX Spark (which looks to be the exact same stack, sans the high capacity network interface) will help this get a bit more attention, but nVidia's long long long war with the open source community is not helping. Focusing on mainlining kernel support would go a long way to getting the community to be supportive.

    Of course, a massive regression just hit Linux 7+/7.1 plus for ROCm hosts, so it's just rough everywhere.

  • cogman10 3 hours ago
    I strongly considered it, but the one thing that scares me away from wanting to do the spark is you basically have to use nvidia's linux (from what I've read) and it doesn't appear the nvidia is interested in upstreaming their kernel changes.

    I'm avoiding where possible buying electronics where support is controlled by the manufacturer and not me.

    • dllu 2 hours ago
      Many other distros almost work out of the box (as in they boot and run without any modifications). The custom kernel patches you mentioned address mainly non-critical bugs such as a bug where the Realtek r8127 stops working after a reboot (but it works if you turn it off and on again) [1] [2]. I'd consider it in a way better state than trying to run other Linux distros on certain device tree-based devices like, say, Qualcomm Snapdragon machines. The regular NVIDIA drivers with the open source kernel modules work just fine. Talos Linux supports DGX Spark since version 1.12 [3]. I also know of people using Fedora and nixOS successfully.

      [1] https://github.com/NVIDIA/NV-Kernels/compare/ea55925ab430f1e...

      [2] https://forums.developer.nvidia.com/t/realtek-r8127-ethernet...

      [3] https://github.com/siderolabs/talos/issues/12170

    • cyril-crutches 2 hours ago
      If I understand your comment correctly, I think he addresses that in the first few paragraphs:

      > The DGX Spark runs “DGX OS” but it is in fact just plain old Ubuntu 24.04 with some additions. If you want, you can just install another Linux distribution easily (Fedora works well), although there may be a couple of weird bugs with the Realtek Ethernet driver so the NVIDIA version of the Linux kernel has a couple of patches. Unlike some other ARM devices, the DGX Spark is all ACPI rather than device tree based, so regular Linux builds for arm64 work just fine.

      • cogman10 32 minutes ago
        Well that is better than what I gleened.

        I thought I'd read that the GPU needed extra kernel patches to properly work. If it's just the Ethernet driver that seems a lot more appealing.

        I thought it was device tree as well, so great that it's actual ACPI.

    • willis936 2 hours ago
      A good instinct. There are a lot of things a $500 AMD GPU can do in linux that a $5000 DGX cannot.
      • girvo 1 hour ago
        Although one of the ones it cannot do is have/address 128GB of video memory, so it depends on what you want to achieve.
      • ARandomerDude 2 hours ago
        I know nothing about this topic but your comment piqued my curiosity. What would a $500 AMD GPU do better than DGX?
        • bigyabai 1 hour ago
          AMD has Mesa drivers for graphics, which are better-optimized than Nvidia's proprietary Linux Vulkan drivers. It can be fixed in software, but Nvidia's only barely started to catch up.

          The focus for Nvidia's GPU stack on Linux is getting CUDA working, which means that some traditional raster features get neglected.

      • colordrops 1 hour ago
        Fair. Honest caveat - I keep seeing "a good instinct" everywhere now. Is this humans acquiring new phrases from Claude? Is there a name for this phenomenon yet?
  • ciupicri 1 hour ago
    Somehow related: "The end of my AArch64 [Ampere Altra Q80-30] desktop experiment", https://news.ycombinator.com/item?id=48728599 / https://marcin.juszkiewicz.com.pl/2026/06/26/the-end-of-the-...
    • dllu 1 hour ago
      The Ampere cpu in that post has a "lack of single core CPU speed" despite ripping through compilation tasks with its 80 cores. But the DGX Spark's CPU has fairly decent single core speed and is more like Apple Silicon in this respect.
  • MrVitaliy 1 hour ago
    I do appreciate how Nvidia tries to say close to vanilla with Linux and Android (nvidia shield). Instead of trying to build a shitty moat like Samsung with all their garbage software.

    If nvidia ever releases Android smartphone, I'd probably stand in line to get one.

    • dietr1ch 1 hour ago
      After leaving a few of great-on-paper SoCs as paperweights I've learnt that I just don't want to deal with anyone's custom platform as I'll eventually be left with an outdated system that's annoying and time-consuming to maintain.
  • jubilee33 1 hour ago
    This is an interesting review. I have a Chinese strix halo box that's isnt available in the west (favm faex1) I've been able to do some ok graphical gen, or some decent agentic tasks as a fallback for when some of the APIs are overloaded during business hours, but nothing amazing for sure, and also not both at the same time. But here's the thing...it cost me 1800usd two months ago....and it's runs x86. I am struggling to see why people pay +2x more for the Arm Nvidia version, despite the slightly higher bandwidth it still does basically the same AI tasks and alot fewer high end general computing tasks... I like my box but I wouldn't find it useful enough to pay more than I did for it or get more of them and cluster for instance. Can anyone explain the allure of the Nvidia box, other than brand name?
    • nightski 1 hour ago
      It's simple, the 395+ Max Strix Halo you bought for $1800 is now a ~$4000 build (at least the AMD AI dev unit). If only we had time travel right? Either way, the Nvidia unit comes with Connect-X 7. That may or may not matter to you, but the hardware for that isn't cheap. In general the Nvidia cards also have better support for models. I know AMD is trying to catch up but anything except their datacenter cards do not seem to be getting a lot of attention.
    • embedding-shape 1 hour ago
      > isnt available in the west [...] Can anyone explain the allure of the Nvidia box

      The first part might answer the second one. Otherwise, the lack of CUDA and the nvidia ecosystem of tooling could also explain why it doesn't seem so interesting for AI tasks.

      • icedchai 1 hour ago
        Yep, stuff is more likely to "just work" on NVidia. Example: pytorch

        In most benchmarks, the Spark is also faster at the prompt processing / prefill phase.

      • jubilee33 1 hour ago
        But strix halo boxes themselves are available, just not that one. And despite my concerns about what's said about Cuda and ROCm I have never had a problem running any model, for image or text or voice, the community has done great work in making things work. So the point of the question stands.

        It's also interesting that the most high end Chinese equipment, both prosumer things like these boxes but also the Huawei professional stack is just not available in the places it would be most appreciated. Not sure if thats china tit for tat, or western "we don't want your commie hardware anyways"

        But for a lot of people it sucks cause nobody should be paying 4.2k for this product. The value isn't there.

        • embedding-shape 1 hour ago
          > I have never had a problem running any model, for image or text or voice, the community has done great work in making things work

          There is a whole world of other tooling and stuff that isn't just for hobbyists to run inference with ML models, but also how to do profiling, debugging and gathering data when you run distributed workloads, and so on. The nsight toolkit seems miles ahead of the competition on other platforms, as just one example.

  • aftbit 2 hours ago
    Neat blog! I was intrigued by this bullet point mentioned in passing:

    >my four hard drive USB 3.2 ZFS raidz2 array with four 24 TB drives

    Can you speak more about this? Which USB array did you choose? How well does it work? I've been slowly planning a transition away from my power-hungry surplus enterprise gear in the 19" rack towards a smaller, quieter, lower power setup ... but storage is the real kicker right now. I have a 12x18TB array in raidz2 built into a 1U NAS case, and I just can't quite figure out a better way to package something like that. I would need three USB arrays if I want to reuse the existing drives, which I think I do given how expensive storage is today.

    • dllu 2 hours ago
      It's an Orico 9948C3 with four Seagate Barracuda 24TB drives. They were on sale last year [1].

      Unfortunately, the enclosure doesn't work super well on Linux. There is a weird bug where the drives don't enumerate when I boot up my computer. This happens on both my x86_64 AMD machine running Linux, and on the DGX Spark. The solution is... simply power cycle the enclosure a couple of times by toggling the power button on it and then it works. Once all four drives show up in lsblk, I can `sudo zfs import ...` manually. This is really gross and annoying. Replacing the USB cable, flipping the USB-C cable 180 degrees, hot plugging it, etc, all didn't work, both on the DGX Spark and the other Linux machine. I've also read reports of it being unstable in UAS mode on Linux but I haven't found a big difference in stability between enabling UAS or falling back to usb-storage.

      Once it starts up correctly though, the drives are fast. I store my huge amount of 100 megapixel photos on it.

      The Seagate Barracudas are helium-filled HAMR/CMR drives and are apparently rebranded/binned Exos drives. They aren't rated for 24/7 use but then neither are the refurbished Exos drives.

      [1] https://www.reddit.com/r/buildapcsales/comments/1p29pm8/hdd_...

      • aftbit 1 hour ago
        Yeah... that's been my past experience with USB docks, at least any with more than one slot. I've never had great luck with them. "Can recover from power outage without being touched" is a key requirement for my NAS so I'll give that one a pass and stick with my "SATA drives directly attached to a SAS controller" strategy for now. Thanks for the reply.

        As for the 24/7 use, yeah so be it. The I in RAID stands for Inexpensive. If they fail after 10 years at 24/7, so be it. I have drive level redundancy and frequent offsite backups of anything critical.

  • bullen 1 hour ago
    I have been running uConsoles with CM5 (2712 and 3588 with 16GB RAM) for 6 months as daily drivers.

    They are ~$500* and present the same ARM problems/opportunities.

    But they are completely silent (no fan, the case is the heat sink).

    My 6600(3050) desktop from 2016(2024) with replaced SSD(2021)/RAM(2025) (they age like milk) now gets little use and M$ will soon sleep with the fishes.

    *Hard to get now as the 3588 that has linux for uConsole is out of stock and the Raspberry one is rare and more expensive by the day.

    • jubilee33 1 hour ago
      Got 2 of these early on. They are wonderfully designed, unfortunately I've been caught in the "building things" trap for few months now and they have been relegated to being used as retro gaming/ computing learning boxes for my sons. They really don't appreciate it at all (yet)....but I can hope they will remember it in the future. Teaching them to get to terminal and run the emulator was great fun...reminded me of MSDoS and the hours of troubleshooting to run games with limited memory and drivers back in the day. I just worry that with LLMs the whole point of teaching them basic terminal/troubleshooting skills might be lost soon. We will see.
  • midnightbobarun 3 days ago
    Super-powerful (if rather pricy) Linux desktop that happens to play games while doing everything else... that man is living my dream :'D
  • skolos 1 hour ago
    "Multi-token prediction gives a free speedup of up to 2x on many models" - at the expense of halving prompt processing speed
    • tingletech 28 minutes ago
      it works quite will for me in llama.cpp, but I get more like 20% to 40% speedup on tokens per second. I generally use

        spec-type = draft-mtp,ngram-mod
        spec-draft-n-max = 4
      
      I have not observed any effect on prompt processing, which is usually an order of magnitude faster than generation on my spark.
    • hedgehog 49 minutes ago
      MTP has no effect on prompt processing.
      • skolos 34 minutes ago
        Interesting - in my setup (llama.cpp rtx5090 qwen-3.6 27b) prompt processing with mtp is almost half vs non mtp. Sounds like I need to investigate what is wrong.
        • hedgehog 18 minutes ago
          Maybe enabling MTP causes some weights to be displaced to host memory? MTP itself doesn't do anything during prefill so that should be exactly unchanged, decode will vary depending on settings but with 2-4 proposals depending on workload I've never seen an overall slowdown.

          edit: I recommend building recent llama.cpp from source, I've been updating about once a week, as there has been a fair amount of work related to MTP recently. If you're running a lot of tool calling on Qwen you might also benefit from one of the bugfixed chat templates like the Froggeric version.

  • biddit 1 hour ago
    Please don’t buy a DGX Spark unless all three of these are true:

      - You value simplicity more than performance or price-to-performance.
      - You accept that the hardware will depreciate rapidly.
      - You’re prepared to buy two or four of them.
    
    OR:

      - You want to run frontier models right now as cheaply as possible
      - You want to run high-parameter models on a 15a breaker/line
    
    Otherwise, get a normal, high-bandwidth GPU.

    A single Spark gives you roughly 115 GB of usable memory compared with the 24–32 GB found on many lower-cost GPUs. It's certainly a big increase, but in practice it does not unlock dramatically better models.

      - One Spark: More memory, but mostly enough for poor-quality, extremely low-bit quants of larger models.  
      - Two Sparks: Enough for mid-tier parameter models at reasonable quants, such as DeepSeek V4 Flash and HY3.
      - Four Sparks: Enough for GLM 5.2 at a reasonable quant.  You'll need a $1000+ switch too.
    
    The problem is that Sparks are slow compared with almost everything else in their price range. Many factors affect inference speed, but memory bandwidth is one of the biggest. A $4,000-plus DGX Spark provides only 273 GB/s.

      | GPU                 | Memory bandwidth |           VRAM | Approx. price  |
      | ------------------- | ---------------: | -------------: | ------------:  |
      | DGX Spark           |         273 GB/s | ~115 GB usable |        $4,000+ |
      | RTX 5060            |         448 GB/s |          16 GB |          $600  |
      | Radeon AI Pro R9700 |         640 GB/s |          32 GB |        $1,200  |
      | RTX 4000 Pro        |         672 GB/s |          24 GB |        $2,300  |
      | RTX 4500 Pro        |         896 GB/s |          32 GB |        $3,500  |
      | RTX 3090            |         936 GB/s |          24 GB |        $1,200  |
      | RTX 5000 Pro        |       1,344 GB/s |          48 GB |        $6,000  |
      | RTX 5090            |       1,792 GB/s |          32 GB |        $4,000  |
      | RTX 6000 Pro        |       1,792 GB/s |          96 GB |       $12,000  |
    
    Yes, the Spark has substantially more memory. But going from roughly 24 GB to 115 GB does not necessarily unlock substantially better model quality. In many cases, it only lets you load heavily compressed 2-bit versions of larger models, such as DeepSeek V4 Flash, with serious quality degradation.

    24–32 GB is currently a sweet spot. Models such as Qwen 3.6 27B and 35B-A3B:

      - Perform far above what their parameter counts suggest.
      - Fit comfortably within 24–32 GB of VRAM at reasonable quantization levels.
    
    A 4-bit quant of Qwen 3.6 27b (18 GB) will out-perform a 2-bit quant of DeepSeek v4 Flash (90gb).

    Instead of the Spark, if I had a roughly $4,000 budget...

    Assuming I already had a reasonably modern desktop:

      - One RTX 5090, RTX 5000 Pro, or RTX 4500 Pro.
      - Two RTX 3090s, RTX 4000 Pros, or R9700s, provided the motherboard can bifurcate two physical x16 slots into x8/x8.
    
    If I were building a system from scratch:

      - A DDR4- or PCIe 4.0-era consumer CPU and motherboard that supports x8/x8 bifurcation.
      - Two RTX 3090s, RTX 4000 Pros, or R9700s.
    
    If I were already planning to buy a new Mac:

      - A MacBook Pro M5 with 64 GB or 128 GB of unified memory.
    
    For context, these are the systems I currently run:

      - EPYC Turin with four RTX 6000 Pro Max-Qs.
      - EPYC Milan with four RTX 3090s.
      - AM4 with two RTX 3090s.
      - AM4 with two RTX 3090s.
      - Intel Raptor Lake with two RTX 5060 Ti.
      - MacBook Pro M3 128GB Unified
    • josh-wrale 40 minutes ago
      I agree with this. I have a dual rtx4090 machine, a 128gb m5 max mbp, and a dgx spark. RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic on the dual rtx4090 machine under vLLM absolutely slays at token speed. I'm frustrated with how hard it is to realize goodness on the dgx spark.
    • pizza234 55 minutes ago
      I'm actually confused by the post - indeed why choosing the 27B dense and the 35B MoE models? Given the structure of the post I honestly have the impression that the Spark was purchased out of curiosity more than for a specific purpose (LLMs).
      • biddit 8 minutes ago
        Ah yeah, your reaction makes sense. In the circles I run in, there is a lot of hype around Sparks for inference, so my gut reaction is to respond with this type of warning.

        I did not intend to imply that the post author was advocating that they're great for inference, as they're obviously not.

      • InTheArena 7 minutes ago
        The 27b dense is _remarkably_ better at coding tasks.

        Especially at larger quanitizations (Q4 is pretty crap).

    • gerdesj 40 minutes ago
      Actually, its ~119GB usable (I have one). You can shutdown a lot of unneeded services if you only use it remote which will trim a lot more fat. You can enable the RDP service if you don't want to sit in front of it but get a desktop interface.

      The "shitty" network is 10Gb/s and wifi7! You get a twin QSFP28DDlol+++ (I jest) that each run at 200Gb/s - not for the casual home user but handy at work, although I "only" have 40Gb/s on my switches sigh. With and no switch two you can do a three node cluster with some careful networking. If you want to do more then a switch is needed and it will need to be pretty funky! That said you could wire them up in a circle and use VLANs and MSTP and accept less than 200Gb/s per link. You'll probably need Openvswitch and a lie down afterwards.

      I'm not a fan of the Gnome desktop but it works well enough and I think the Nvidia customised Ubuntu is well thought out. You get all the complicated NVidia extras pre-installed, along with docker (full fat, not the Ubuntu one) for a fairly quick start. It includes Ubuntu Pro which is free for five systems anyway but its nice to see it pre-installed.

      You can run quite decent models on this thing see: https://spark-arena.com/ Also see "DS4".

      We blew abut £4000 on one and it will pay for itself in a few months. I tried pricing up an Apple thingie and the Store wouldn't offer me more than 96Gb of RAM and a delivery date in Q3 at the earliest. Our Spark rocked up next day. They seem to come in 1TB or 4TB SSD variants. 1TB is enough for me and saves a lot of cash - keep an eye on your model downloads and ruthlessly delete old experiments. docker system prune.

      We went for the Asus variant that has active cooling and I stuck it in the ceiling cable tray over our computer room racks. It sits on 1½" stainless steel mesh with lots of clearance in an actively cooled environment.

  • zer0zzz 17 minutes ago
    This is cool. Can I run the same fedora as the asahi Linux I have on my M2 Ultra and use the M2 to do builds and the DGX to launch kernels? That'd be the ideal fully arm at-home coding setup I think.
  • ShipVoicedev 52 minutes ago
    Is Nvidia better than Intel
  • ramshanker 2 hours ago
    Yes. Waiting for the Windows Version myseflf. RTX Spark Desktop.
    • dijit 2 hours ago
      I guess copilot needs all the help it can get?
  • bryanlarsen 3 hours ago
    It's interesting how many of these issues don't appear to be specific to the DGX Spark but to the standard "Nvidia GPUs suck on Linux" type of issues that afflict a lot of people.
    • a-dub 3 hours ago
      it's actually very stable on x86_64 these days, even with optimus.
  • haunter 2 hours ago
    This is something I'd do if I've had the disposable income lol

    >Non-Steam games have a lower chance of working

    Wonder if it's true for GOG games because they are usually installed in a neatly packaged folder without any bloat.

  • rvz 2 hours ago
    I would avoid the DGX Spark. For that price and its performance on running local models it is a complete scam. This tweet says it all [0]

    [0] https://xcancel.com/petergostev/status/1978230978725507108

  • trentor 2 hours ago
    I'm genuinely disappointed with my Spark. I don't know how anyone can claim it performs decently with LLMs or diffusion models. Back when I worked in VFX in the early 2000s, we had a saying: "Render time is coffee time" and if you try to run this thing with a usable context size, you'll be drinking a lot of coffee. Most of the optimizations it relies on for inference simply aren't available for training, so it crawls like a snail on almost every model. An RTX 6000 Blackwell would have been the better investment for an AI enthusiasts and for general computing there are cheaper offerings.
    • mapontosevenths 2 hours ago
      If you bought it for inference you made a mistake. They aren't good at that. Use it to train models and experiment with ML. It's much better at that.

      If you just want local inference buy a Mac.

      If you bought early on, like I did, the Spark is probably worth double what you payed now. I think I paid $3,000 retail for mine and the last time I looked they were fetching close to $6k on ebay. I'm not sure if that's still the case, but you can buy a very nice Mac with $6k.

      • trentor 2 hours ago
        Mhm... maybe read my full comment?
        • mapontosevenths 2 hours ago
          Are you saying that it's slow for training? Sorry, your comment is confusingly worded to me.

          I've not had any issues in that regard, but I'm working with LLM's not training diffusion models. Are you following one of the Nvidia provided recipes or inventing something on your own? The last time I looked into it they benchmarked very well, but we both know that doesn't always mean much.

    • dllu 2 hours ago
      The DGX Spark hits a sweet spot for me where it can simultaneously work for general computing and run local LLM inference fast enough for some hobbyist dabbling. The RTX 6000 Pro Blackwell is more than twice the price (it has increased quite a bit recently from $8000 to $11600), not to mention the "rest of the PC" needed to get it working, so it's not really a fair comparison. Compared to other 128 GB unified memory devices like the Mac Studio and the Strix Halo, the DGX Spark fairly priced in my opinion.
      • icedchai 1 hour ago
        I have a AMD Strix Halo box I use for similar dabbling. It definitely wasn't an "out of the box" experience, fiddling around with kernel versions and ROCm installs. These days I mostly wind up using the Vulkan build of llama.cpp for inference.
    • Foobar8568 1 hour ago
      I have access to both a RTX 5090 PC with 64gb of ram and a Spark 128gb, the performance of the Spark has been highly disappointing.

      I prefer 1000x the RTX one, even with 64gb of ram.