ClusterUpdated 2026-07-03 · 12 min read · by RTXsparks Lab
2-Node RTX Spark Cluster: networking (2026)
Networking for a 2-node RTX Spark cluster: 10GbE vs 25GbE vs InfiniBand, switch choice, and NCCL tuning.
Bill of materials
2 × Spark 128GB nodes, 1 × Arista 7050X 32×25GbE switch, ConnectX-6 25GbE NICs, DAC cables, and a 42U rack.
| Item | Qty | Unit $ | Total $ |
|---|---|---|---|
| Spark node | 2 | 4,299 | 8,598 |
| 25GbE switch | 1 | 6,800 | 6,800 |
| NIC + DAC | 2 | 420 | 840 |
Networking
NCCL over RoCE v2 on 25GbE gives ~2.9 GB/s per pair; sufficient for TP=2 on 70B and TP=2 + PP=2 on 200B.
Orchestration
Ray Serve for inference, Slurm for training. Nodes come up via PXE, driver stack installed by cloud-init.
Model placement
TP=2 for 70B, TP=1 × DP=2 for parallel serving, or run one 235B MoE with expert parallelism 2-way.
Cost & power
Full cluster draws 340W sustained; at $0.16/kWh, that's $39/month.
Frequently asked questions
Is 2 nodes enough for Llama 405B?
Not comfortably — go 6 or 8 nodes.
Do I need InfiniBand?
No, 25GbE is fine at this scale.