ClusterUpdated 2026-07-03 · 12 min read · by RTXsparks Lab
6-Node RTX Spark Cluster: scaling to large models (2026)
Running 200B+ models on a 6-node RTX Spark cluster with FSDP and vLLM tensor parallel.
Bill of materials
6 × Spark 128GB nodes, 1 × Arista 7050X 32×25GbE switch, ConnectX-6 25GbE NICs, DAC cables, and a 42U rack.
| Item | Qty | Unit $ | Total $ |
|---|---|---|---|
| Spark node | 6 | 4,299 | 25,794 |
| 25GbE switch | 1 | 6,800 | 6,800 |
| NIC + DAC | 6 | 420 | 2,520 |
Networking
NCCL over RoCE v2 on 25GbE gives ~2.9 GB/s per pair; sufficient for TP=6 on 70B and TP=6 + PP=2 on 200B.
Orchestration
Ray Serve for inference, Slurm for training. Nodes come up via PXE, driver stack installed by cloud-init.
Model placement
TP=6 for 70B, TP=3 × DP=2 for parallel serving, or run one 235B MoE with expert parallelism 6-way.
Cost & power
Full cluster draws 1020W sustained; at $0.16/kWh, that's $118/month.
Frequently asked questions
Is 6 nodes enough for Llama 405B?
Yes, at Q3_K_M with TP=6.
Do I need InfiniBand?
Recommended for training; inference works on 25GbE.