How-toUpdated 2026-07-01 · 9 min read · by RTXsparks Lab
Serving Mistral Large 2411 with Ollama on RTX Spark
Production Ollama deployment of Mistral Large 2411 on RTX Spark. Configs, systemd, TLS, observability.
Install
Install Ollama via pip/docker/apt as appropriate. Verify GPU visibility with nvidia-smi.
- # Ollama install
- curl -fsSL https://ollama.com/install.sh | sh
Config
Optimized Ollama config for Mistral Large 2411 at Q4_K_M: batch, KV cache dtype, and parallelism knobs.
Systemd unit
[Unit] Description=Ollama for Mistral Large 2411 — [Service] ExecStart=/usr/bin/ollama serve mistral-large-2411 --port 8000 — Restart=always.
nginx + TLS
Terminate TLS at nginx with a Let's Encrypt cert, rate-limit at 20 req/s per IP, and proxy to the local Ollama port.
Observability
Scrape Ollama Prometheus metrics into Grafana; watch tokens_per_second, batch_size, and kv_cache_usage.
Frequently asked questions
Is Ollama the fastest for Mistral Large 2411?
It's the easiest, not always the fastest.
Can I run multiple models with one instance?
Yes, it hot-swaps.