ClusterUpdated 2026-07-03 · 12 min read · by RTXsparks Lab

8-Node RTX Spark Cluster: scaling to large models (2026)

Running 200B+ models on a 8-node RTX Spark cluster with FSDP and vLLM tensor parallel.

Bill of materials

8 × Spark 128GB nodes, 1 × Arista 7050X 32×25GbE switch, ConnectX-6 25GbE NICs, DAC cables, and a 42U rack.

ItemQtyUnit $Total $
Spark node84,29934,392
25GbE switch16,8006,800
NIC + DAC84203,360

Networking

NCCL over RoCE v2 on 25GbE gives ~2.9 GB/s per pair; sufficient for TP=8 on 70B and TP=8 + PP=2 on 200B.

Orchestration

Ray Serve for inference, Slurm for training. Nodes come up via PXE, driver stack installed by cloud-init.

Model placement

TP=8 for 70B, TP=4 × DP=2 for parallel serving, or run one 235B MoE with expert parallelism 8-way.

Cost & power

Full cluster draws 1360W sustained; at $0.16/kWh, that's $157/month.

Frequently asked questions

Is 8 nodes enough for Llama 405B?

Yes, at Q3_K_M with TP=8.

Do I need InfiniBand?

Recommended for training; inference works on 25GbE.

Related guides