INSTALLATION

One command, every environment.

Install SLM on macOS, Linux, or a CUDA server. The engine is the same everywhere; only the model backend changes.

01 · Prerequisites

Prepare the environment

# Python 3.11+ and Git are the only hard requirements
$ python3 --version
Python 3.11.9

# Optional but recommended: a local model server
# macOS:   brew install ollama        (or use MLX)
# Linux:   curl -fsSL https://ollama.com/install.sh | sh
$ ollama pull llama3.1:8b

02 · One-command setup

Clone, install, and go

$ git clone <your-slm-repo> && cd slm
$ python3 -m venv .venv
$ .venv/bin/pip install -r requirements.txt

# Start the web UI
$ make ui        # -> http://localhost:8601

# Train a pack and serve it (OpenAI-compatible)
$ make train PACK=my-corpus
$ make serve PACK=my-corpus PORT=8081
The make train target does everything: ingest, structure, datagen, training, evaluation, gating, and adapter registration. Hyperparameters are chosen automatically from corpus size and available memory.

03 · By environment

Platform-specific installation

 macOS (Apple Silicon)

Use MLX for fast local inference and LoRA training, or ollama for the simplest setup.

brew install ollama
ollama pull llama3.1:8b
python3.11 -m venv .venv
.venv/bin/pip install -r requirements.txt
make ui

🖥 Linux (CPU)

CPU-only works fine for grounded answering; the retrieval index runs on CPU regardless.

curl -fsSL https://ollama.com/install.sh | sh
ollama pull llama3.1:8b
sudo apt install -y python3.11-venv
python3.11 -m venv .venv
.venv/bin/pip install -r requirements.txt
make ui

⚡ Linux + CUDA GPU

For production scale, run vLLM (OpenAI-compatible, multi-adapter LoRA) with PEFT/TRL for training.

pip install -r requirements-cuda.txt
# vLLM server with LoRA adapters
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --enable-lora --lora-modules pack1=/adapters/1
# training: PEFT/TRL (rank-64 QLoRA + DPO)
python -m engine.train.pipeline PACK=my-corpus

🛡 Air-gapped / offline

Pull models once on a connected machine, cache them, then run fully offline with a local judge.

# on the connected machine:
ollama pull llama3.1:8b
ollama pull qwen2.5:3b   # local judge
scp -r ~/.ollama offline-host:~/.ollama

# on the offline host:
export OLLAMA_HOST=127.0.0.1:11434
export JUDGE_PROVIDER=ollama
make ui

04 · Supported agents & LLMs

What we support and how to connect it

Every pack is exposed as an OpenAI-compatible endpoint, so any agent, copilot, or tool that speaks OpenAI's API can use it. Point your client at the pack's URL and it becomes a domain specialist.

Agents & integrations

  • Any OpenAI-compatible client (curl, Python, LangChain, LlamaIndex, AutoGen)
  • Custom copilots and RAG pipelines pointing at /v1/chat/completions
  • Slack/Discord bots via a 10-line webhook bridge
  • CLI usage via engine.serve.api / engine.serve.ollama

LLM backends

  • ollama — llama3.1, qwen, mistral, gemma, etc.
  • vLLM — any HF model, LoRA adapters, CUDA
  • MLX — Apple Silicon models
  • transformers + PEFT — CPU/GPU, adapter loading
  • Judge — local (qwen2.5:3b) or any OpenAI-compatible API
# Connect any OpenAI-compatible agent to a pack
$ curl http://localhost:8081/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{"model":"my-corpus",
         "messages":[{"role":"user","content":"What does the policy say?"}]}'
# -> {"choices":[{"message":{"content":"Per Section 4.2 ..."}}]}