BenchmarkUpdated 2026-07-01 · 12 min read · by RTXsparks Lab
vLLM Performance Tuning on RTX Spark
Squeeze maximum throughput from vLLM on RTX Spark: chunked prefill, speculative decoding, NVFP4, and paged-attention tuning.
Config
See sample YAML in docs repo.
Frequently asked questions
Batch size?
Start at 8; scale until GPU util >92%.