DEPLOYMENT

Runs where your data lives.

SLM is a self-hosted engine. Deploy it on a laptop, a single VM, or an air-gapped cluster — the same one-command workflow applies. Your documents stay on your infrastructure.

Architectures

One engine, three deployment postures

📱

Workstation

Run the full pipeline on a laptop or desktop. Apple Silicon (MLX) and Linux (CPU) supported. Ideal for a single user or a team prototype.

macOS · Linux · 16–128 GB RAM

🖥

Single VM / server

Deploy as systemd services (or Docker) on one machine with nginx in front. Multi-pack, login, and per-pack access control. This is how we run it in production.

Linux · systemd · nginx · TLS

🛡

Air-gapped / restricted

No external calls required. Models are pulled once and cached; the judge can run locally. Built for restricted and offline networks.

offline · local judge · local models

Reference stack

The production stack we run

Our reference deployment (this site's parent service) runs on a GCP VM with:

  • Model backend: ollama (local llama3.1:8b) or vLLM on CUDA
  • Web UI + API: FastAPI behind nginx with TLS (Let’s Encrypt)
  • Retrieval: hybrid BM25 + dense (sentence-transformers) + FAISS
  • Process supervision: systemd user services with auto-restart
  • Access control: IP allow-list + per-pack permissions + iptables
  • Backups: daily archive of adapters, configs, and indexes
# deploy/ui: systemd user services
$ systemctl --user status expertpacks-ui
  active · UI on :8088 (nginx + TLS)

# deploy/model: ollama or vLLM
$ curl localhost:11434/api/tags
  llama3.1:8b · qwen2.5:3b (judge)

# reverse proxy (nginx)
server_name slm.qaso.ai;
proxy_pass http://127.0.0.1:8088;

Backends

Model backends we support

BackendWhereUse caseNotes
ollamaAny Linux/macOSEasiest local/VM model serverllama3.1, qwen, mistral, many more
vLLMLinux + CUDA GPUProduction scale, multi-adapterOpenAI-compatible, LoRA adapters
MLXApple SiliconLocal Mac inference + LoRA trainingmlx-lm
transformers + PEFTAny CPU/GPUCPU serving, adapter loadingHuggingFace stack
The engine is model-agnostic: the same packs and UI work across all four backends. Point the engine at a new backend by changing the serve command — no code changes.