ClusterUpdated 2026-07-03 · 12 min read · by RTXsparks Lab
8-Node RTX Spark Cluster: cost analysis (2026)
3-year TCO for a 8-node RTX Spark cluster including capex, power, and vs cloud.
Bill of materials
8 × Spark 128GB nodes, 1 × Arista 7050X 32×25GbE switch, ConnectX-6 25GbE NICs, DAC cables, and a 42U rack.
| Item | Qty | Unit $ | Total $ |
|---|---|---|---|
| Spark node | 8 | 4,299 | 34,392 |
| 25GbE switch | 1 | 6,800 | 6,800 |
| NIC + DAC | 8 | 420 | 3,360 |
Networking
NCCL over RoCE v2 on 25GbE gives ~2.9 GB/s per pair; sufficient for TP=8 on 70B and TP=8 + PP=2 on 200B.
Orchestration
Ray Serve for inference, Slurm for training. Nodes come up via PXE, driver stack installed by cloud-init.
Model placement
TP=8 for 70B, TP=4 × DP=2 for parallel serving, or run one 235B MoE with expert parallelism 8-way.
Cost & power
Full cluster draws 1360W sustained; at $0.16/kWh, that's $157/month.
Frequently asked questions
Is 8 nodes enough for Llama 405B?
Yes, at Q3_K_M with TP=8.
Do I need InfiniBand?
Recommended for training; inference works on 25GbE.