Install SLM on macOS, Linux, or a CUDA server. The engine is the same everywhere; only the model backend changes.
01 · Prerequisites
# Python 3.11+ and Git are the only hard requirements $ python3 --version Python 3.11.9 # Optional but recommended: a local model server # macOS: brew install ollama (or use MLX) # Linux: curl -fsSL https://ollama.com/install.sh | sh $ ollama pull llama3.1:8b
02 · One-command setup
$ git clone <your-slm-repo> && cd slm $ python3 -m venv .venv $ .venv/bin/pip install -r requirements.txt # Start the web UI $ make ui # -> http://localhost:8601 # Train a pack and serve it (OpenAI-compatible) $ make train PACK=my-corpus $ make serve PACK=my-corpus PORT=8081
make train target does everything: ingest, structure, datagen, training, evaluation, gating, and adapter registration. Hyperparameters are chosen automatically from corpus size and available memory.03 · By environment
Use MLX for fast local inference and LoRA training, or ollama for the simplest setup.
brew install ollama ollama pull llama3.1:8b python3.11 -m venv .venv .venv/bin/pip install -r requirements.txt make ui
CPU-only works fine for grounded answering; the retrieval index runs on CPU regardless.
curl -fsSL https://ollama.com/install.sh | sh ollama pull llama3.1:8b sudo apt install -y python3.11-venv python3.11 -m venv .venv .venv/bin/pip install -r requirements.txt make ui
For production scale, run vLLM (OpenAI-compatible, multi-adapter LoRA) with PEFT/TRL for training.
pip install -r requirements-cuda.txt # vLLM server with LoRA adapters vllm serve meta-llama/Llama-3.1-8B-Instruct \ --enable-lora --lora-modules pack1=/adapters/1 # training: PEFT/TRL (rank-64 QLoRA + DPO) python -m engine.train.pipeline PACK=my-corpus
Pull models once on a connected machine, cache them, then run fully offline with a local judge.
# on the connected machine: ollama pull llama3.1:8b ollama pull qwen2.5:3b # local judge scp -r ~/.ollama offline-host:~/.ollama # on the offline host: export OLLAMA_HOST=127.0.0.1:11434 export JUDGE_PROVIDER=ollama make ui
04 · Supported agents & LLMs
Every pack is exposed as an OpenAI-compatible endpoint, so any agent, copilot, or tool that speaks OpenAI's API can use it. Point your client at the pack's URL and it becomes a domain specialist.
/v1/chat/completionsengine.serve.api / engine.serve.ollama# Connect any OpenAI-compatible agent to a pack $ curl http://localhost:8081/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{"model":"my-corpus", "messages":[{"role":"user","content":"What does the policy say?"}]}' # -> {"choices":[{"message":{"content":"Per Section 4.2 ..."}}]}