BenchmarkUpdated 2026-07-01 · 12 min read · by RTXsparks Lab

vLLM Performance Tuning on RTX Spark

Squeeze maximum throughput from vLLM on RTX Spark: chunked prefill, speculative decoding, NVFP4, and paged-attention tuning.

Config

See sample YAML in docs repo.

Frequently asked questions

Batch size?

Start at 8; scale until GPU util >92%.

Related guides