Documentation
Everything we support, and how to use it. This page is the canonical usage guide.
From a clean machine, following only this page, you can reach a working chat UI in under 30 minutes.
# 1. Clone and install $ git clone <your-slm-repo> && cd slm $ python3.11 -m venv .venv $ .venv/bin/pip install -r requirements.txt # 2. (Recommended) start a local model server $ ollama pull llama3.1:8b # 3. Launch the web UI $ make ui # -> open http://localhost:8601 # 4. Train a pack from a document and chat $ make train PACK=my-corpus # ingest + datagen + train + register $ make serve PACK=my-corpus PORT=8081
A pack is one corpus plus its configuration. It is a folder with a pack.yaml manifest. The engine contains zero corpus-specific code — all specifics live in the pack.
packs/<name>/ pack.yaml # title, sources, chunking, system prompt, eval settings raw/ # original documents processed/ # structured, cleaned text (generated) data/ # training data + manifest (generated) adapters/ # versioned model adapters (generated) reports/ # corpus report, eval report, reflexion logs
To add a pack: create the folder, drop documents in raw/, write a pack.yaml (see any example), then run make train PACK=<name>.
# packs/example/pack.yaml
name: example
title: My knowledge base
sources:
- id: main
file: handbook.pdf
kind: pdf
chunking:
mode: article
eval:
judge_model: qwen2.5:3b
training:
base_model: llama3.1:8b
The web UI accepts drag-and-drop uploads of PDF, Word (.docx), Excel (.xlsx), PowerPoint (.pptx), Markdown, CSV, and plain text, plus URLs. Upload, click Train, and watch plain-language progress.
Each pack has its own chat window. Every answer is retrieval-grounded: the engine retrieves the relevant passages, generates from them, verifies each claim, and returns citations.
[UNSUPPORTED].See the Install page for platform-specific instructions (macOS, Linux CPU, CUDA GPU, air-gapped). Requirements: Python 3.11+, and a model backend (ollama, vLLM, MLX, or transformers).
SLM runs on four supported environments. The retrieval, grounding, UI, and evaluation layers are identical everywhere; only the model backend differs.
| Environment | Backend | Best for | GPU |
|---|---|---|---|
| macOS (Apple Silicon) | MLX or ollama | Local single-user, LoRA training | Apple GPU (unified) |
| Linux (CPU) | ollama / transformers | Self-hosted VM, no GPU | None needed |
| Linux + CUDA | vLLM + PEFT/TRL | Production, multi-adapter, DPO | NVIDIA GPU |
| Air-gapped / offline | ollama (local) | Restricted networks | Optional |
Every pack is exposed as an OpenAI-compatible endpoint (/v1/chat/completions, /v1/models, /health). Any agent that speaks OpenAI's API can connect:
curl http://localhost:8081/v1/chat/completionsThe engine is model-agnostic. Verified backends and model families:
| Backend | Model families | Notes |
|---|---|---|
ollama | Llama 3.x, Qwen 2.5/3, Mistral, Gemma, DeepSeek | Simplest local server; pull any GGUF |
vLLM | Any HuggingFace model | OpenAI-compatible; LoRA multi-adapter |
MLX | Llama, Qwen, Mistral (MLX format) | Apple Silicon; LoRA training built in |
transformers+PEFT | Any HF CausalLM | CPU serving; adapter loading |
Judge | qwen2.5:3b (local) or any API | Used for evaluation only |
# OpenAI-compatible server for a pack (ollama backend) $ python -m engine.serve.ollama my-corpus --port 8081 --model llama3.1:8b # MLX backend (Apple) $ make serve PACK=my-corpus PORT=8081 # Health check $ curl localhost:8081/health {"status":"ok","model":"my-corpus","loaded":true}
Every pack gets a closed-book evaluation on held-out sections, graded for correctness, hallucination rate, correct abstention, and latency. Results are written to packs/<name>/reports/ and surfaced as a quality grade in the UI.
$ make eval PACK=my-corpus | variant | accuracy | halluc. | abstention | latency | | base | 0.00 | 0.13 | - | 0.5s | | trained | 1.00 | 0.00 | 0.75 | 1.6s |
make ui (UI), make serve PACK=x PORT=8081 (model)pack.yaml, put docs in raw/, make train PACK=xmake setup (re-ingest), make datagen, make trainadapter_registry.json at the previous lora_N, restart the serverpacks/*/{pack.yaml,adapters,data,processed} daily/admin/<pack> shows logs, eval tables, adapter versions