What It Is
The DGX Spark is NVIDIA's answer to "I want the big models but I don't want to build a machine or rent a cloud." A GB10 chip with 128 GB of LPDDR5X unified memory and 1 PFLOPS of FP4 compute sits in a 1.13-liter box that plugs into a wall outlet and talks to the world over a 200 Gb fabric port. Software is CUDA-everywhere: Ollama, llama.cpp, vLLM, PyTorch — the same stack the data center runs, no porting.
💡 The 2026 inversion: with GPUs 50–150% over MSRP, a 32 GB DIY build costs more than the Spark while holding a quarter of the model. The Spark now undercuts the 5090 desktop on price and beats it 4× on model size — the trade is raw speed (see below).
Specs & Price (August 2026)
| Spec | Value |
|---|---|
| Chip | NVIDIA GB10 (Grace Blackwell) — 1 PFLOPS FP4 |
| Memory | 128 GB unified LPDDR5X |
| Fabric | 200 Gb ConnectX — dual-Spark clustering supported |
| Price | $4,699 official (retail $4,000–5,615) |
| Power | ~40–45 W idle → ~25 W after the Feb 2026 update; ~135–145 W from the wall during LLM runs (240 W PSU) |
| Noise / form | quiet (minor coil whine reported) · 1.13 L · 4/10 repairability (est.) — 128 GB soldered; 2242 SSD user-swappable |
What It Runs (newest models only)
| Model | Speed |
|---|---|
| gpt-oss-120B Q4 | 4–6 tok/s |
| Mistral Medium 3 (Jul 2026) | 6–8 tok/s |
| MiniMax M2.5 (230B, Q4 ≈ 130 GB) | right at the edge — Q3 or API instead |
| GLM-4.7-Flash (30B-A3B) | 12–15 tok/s |
| Gemma 3 27B (Jul 2026) | ~15 tok/s |
| Ministral 3 14B · Devstral Small 2 | 20–40 tok/s |
Read it honestly: the Spark is a capacity machine. A 5090 desktop is ~4× faster on models that fit 32 GB; the Spark wins only when the model is bigger than 32 GB — which is exactly the tier nothing else under $5K touches.
The Economics
amortized over 5 years — the worst cost-per-token of the picks, because it's bought for capacity, not economy
efficient per watt, but the slow decode caps throughput
Buy It If / Skip It If
Buy it if: you want the biggest newest models on your desk — up to 120B — with zero building, CUDA-everywhere software, and a machine that's quiet and always on.
Skip it if: you want max speed on 27–30B models, gaming, or image generation — those are the 5090 tower's strengths. Or if your needs stop at ≤14B — the $1,465 tower is a fifth of the price.