SLM is a self-hosted engine. Deploy it on a laptop, a single VM, or an air-gapped cluster — the same one-command workflow applies. Your documents stay on your infrastructure.
Architectures
Run the full pipeline on a laptop or desktop. Apple Silicon (MLX) and Linux (CPU) supported. Ideal for a single user or a team prototype.
macOS · Linux · 16–128 GB RAM
Deploy as systemd services (or Docker) on one machine with nginx in front. Multi-pack, login, and per-pack access control. This is how we run it in production.
Linux · systemd · nginx · TLS
No external calls required. Models are pulled once and cached; the judge can run locally. Built for restricted and offline networks.
offline · local judge · local models
Reference stack
Our reference deployment (this site's parent service) runs on a GCP VM with:
# deploy/ui: systemd user services $ systemctl --user status expertpacks-ui active · UI on :8088 (nginx + TLS) # deploy/model: ollama or vLLM $ curl localhost:11434/api/tags llama3.1:8b · qwen2.5:3b (judge) # reverse proxy (nginx) server_name slm.qaso.ai; proxy_pass http://127.0.0.1:8088;
Backends
| Backend | Where | Use case | Notes |
|---|---|---|---|
ollama | Any Linux/macOS | Easiest local/VM model server | llama3.1, qwen, mistral, many more |
vLLM | Linux + CUDA GPU | Production scale, multi-adapter | OpenAI-compatible, LoRA adapters |
MLX | Apple Silicon | Local Mac inference + LoRA training | mlx-lm |
transformers + PEFT | Any CPU/GPU | CPU serving, adapter loading | HuggingFace stack |