The $4,700 AI Supercomputer That Sits on Your Desk

Published: August 9, 2026 — NVIDIA's DGX Spark packs 128 GB of unified memory and 1 PFLOPS into a 1.13-liter box for $4,699. It runs the biggest newest open models — gpt-oss-120B, Mistral Medium 3, MiniMax M2.5 — locally, on your desk, with zero building. In the 2026 shortage market, it quietly became the best capacity-per-dollar buy under $5K.

💼 Quick Takeaways

What It Is

The DGX Spark is NVIDIA's answer to "I want the big models but I don't want to build a machine or rent a cloud." A GB10 chip with 128 GB of LPDDR5X unified memory and 1 PFLOPS of FP4 compute sits in a 1.13-liter box that plugs into a wall outlet and talks to the world over a 200 Gb fabric port. Software is CUDA-everywhere: Ollama, llama.cpp, vLLM, PyTorch — the same stack the data center runs, no porting.

💡 The 2026 inversion: with GPUs 50–150% over MSRP, a 32 GB DIY build costs more than the Spark while holding a quarter of the model. The Spark now undercuts the 5090 desktop on price and beats it 4× on model size — the trade is raw speed (see below).

Specs & Price (August 2026)

SpecValue
ChipNVIDIA GB10 (Grace Blackwell) — 1 PFLOPS FP4
Memory128 GB unified LPDDR5X
Fabric200 Gb ConnectX — dual-Spark clustering supported
Price$4,699 official (retail $4,000–5,615)
Power~40–45 W idle → ~25 W after the Feb 2026 update; ~135–145 W from the wall during LLM runs (240 W PSU)
Noise / formquiet (minor coil whine reported) · 1.13 L · 4/10 repairability (est.) — 128 GB soldered; 2242 SSD user-swappable

What It Runs (newest models only)

ModelSpeed
gpt-oss-120B Q44–6 tok/s
Mistral Medium 3 (Jul 2026)6–8 tok/s
MiniMax M2.5 (230B, Q4 ≈ 130 GB)right at the edge — Q3 or API instead
GLM-4.7-Flash (30B-A3B)12–15 tok/s
Gemma 3 27B (Jul 2026)~15 tok/s
Ministral 3 14B · Devstral Small 220–40 tok/s

Read it honestly: the Spark is a capacity machine. A 5090 desktop is ~4× faster on models that fit 32 GB; the Spark wins only when the model is bigger than 32 GB — which is exactly the tier nothing else under $5K touches.

The Economics

~$24 per million tokens
amortized over 5 years — the worst cost-per-token of the picks, because it's bought for capacity, not economy
~0.36M tokens/kWh
efficient per watt, but the slow decode caps throughput
vs API: DeepSeek V4 Flash (~$0.24/day) is cheaper than the Spark's amortized cost — local only wins here for privacy and fixed bills
vs 5090: $4,699 vs ~$5,660, and the Spark holds gpt-oss-120B — 4× the model, half the speed

Buy It If / Skip It If

Buy it if: you want the biggest newest models on your desk — up to 120B — with zero building, CUDA-everywhere software, and a machine that's quiet and always on.

Skip it if: you want max speed on 27–30B models, gaming, or image generation — those are the 5090 tower's strengths. Or if your needs stop at ≤14B — the $1,465 tower is a fifth of the price.