Serving Llama 3.2 11B Vision with llama.cpp on RTX Spark
Production llama.cpp deployment of Llama 3.2 11B Vision on RTX Spark. Configs, systemd, TLS, observability.
Install
Install llama.cpp via pip/docker/apt as appropriate. Verify GPU visibility with nvidia-smi.
- # llama.cpp install
- brew install llama.cpp
Config
Optimized llama.cpp config for Llama 3.2 11B Vision at Q5_K_M: batch, KV cache dtype, and parallelism knobs.
Systemd unit
[Unit] Description=llama.cpp for Llama 3.2 11B Vision — [Service] ExecStart=/usr/bin/llama-cpp serve llama-3-2-11b-vision --port 8000 — Restart=always.
nginx + TLS
Terminate TLS at nginx with a Let's Encrypt cert, rate-limit at 20 req/s per IP, and proxy to the local llama.cpp port.
Observability
Scrape llama.cpp Prometheus metrics into Grafana; watch tokens_per_second, batch_size, and kv_cache_usage.
Frequently asked questions
Is llama.cpp the fastest for Llama 3.2 11B Vision?
It's the easiest, not always the fastest.
Can I run multiple models with one instance?
One process per model is standard.