ClusterUpdated 2026-07-03 · 12 min read · by RTXsparks Lab
3-Node RTX Spark Cluster: orchestration (2026)
Orchestrating a 3-node RTX Spark cluster with Ray, K8s, or Slurm.
Bill of materials
3 × Spark 128GB nodes, 1 × Arista 7050X 32×25GbE switch, ConnectX-6 25GbE NICs, DAC cables, and a 42U rack.
| Item | Qty | Unit $ | Total $ |
|---|---|---|---|
| Spark node | 3 | 4,299 | 12,897 |
| 25GbE switch | 1 | 6,800 | 6,800 |
| NIC + DAC | 3 | 420 | 1,260 |
Networking
NCCL over RoCE v2 on 25GbE gives ~2.9 GB/s per pair; sufficient for TP=3 on 70B and TP=3 + PP=2 on 200B.
Orchestration
Ray Serve for inference, Slurm for training. Nodes come up via PXE, driver stack installed by cloud-init.
Model placement
TP=3 for 70B, TP=1.5 × DP=2 for parallel serving, or run one 235B MoE with expert parallelism 3-way.
Cost & power
Full cluster draws 510W sustained; at $0.16/kWh, that's $59/month.
Frequently asked questions
Is 3 nodes enough for Llama 405B?
Not comfortably — go 6 or 8 nodes.
Do I need InfiniBand?
No, 25GbE is fine at this scale.