Documentation

How to use SLM

Everything we support, and how to use it. This page is the canonical usage guide.

Quickstart

From a clean machine, following only this page, you can reach a working chat UI in under 30 minutes.

# 1. Clone and install
$ git clone <your-slm-repo> && cd slm
$ python3.11 -m venv .venv
$ .venv/bin/pip install -r requirements.txt

# 2. (Recommended) start a local model server
$ ollama pull llama3.1:8b

# 3. Launch the web UI
$ make ui
# -> open http://localhost:8601

# 4. Train a pack from a document and chat
$ make train PACK=my-corpus   # ingest + datagen + train + register
$ make serve PACK=my-corpus PORT=8081

Packs

A pack is one corpus plus its configuration. It is a folder with a pack.yaml manifest. The engine contains zero corpus-specific code — all specifics live in the pack.

packs/<name>/
  pack.yaml      # title, sources, chunking, system prompt, eval settings
  raw/           # original documents
  processed/     # structured, cleaned text (generated)
  data/          # training data + manifest (generated)
  adapters/      # versioned model adapters (generated)
  reports/       # corpus report, eval report, reflexion logs

To add a pack: create the folder, drop documents in raw/, write a pack.yaml (see any example), then run make train PACK=<name>.

# packs/example/pack.yaml
name: example
title: My knowledge base
sources:
  - id: main
    file: handbook.pdf
    kind: pdf
chunking:
  mode: article
eval:
  judge_model: qwen2.5:3b
training:
  base_model: llama3.1:8b

Uploading documents

The web UI accepts drag-and-drop uploads of PDF, Word (.docx), Excel (.xlsx), PowerPoint (.pptx), Markdown, CSV, and plain text, plus URLs. Upload, click Train, and watch plain-language progress.

  • PDFs are extracted with structure (chapters, sections, footnotes) and cleaned of headers/footers.
  • Text-based formats are ingested as paragraphs (article mode).
  • Word/Excel/PPT are currently handled as text; dedicated binary extraction is on the roadmap.
  • Uploaded files are registered automatically and re-indexed on demand.

Chat & citations

Each pack has its own chat window. Every answer is retrieval-grounded: the engine retrieves the relevant passages, generates from them, verifies each claim, and returns citations.

  • Citations — click any source to open the passage with surrounding context.
  • Strict grounding toggle — on: unsupported claims are dropped; off: they are marked [UNSUPPORTED].
  • Sources only toggle — returns the raw supporting passages instead of prose.
  • Conversation history — maintained per pack.
  • Debug view — retrieved chunks with scores.

Installation

See the Install page for platform-specific instructions (macOS, Linux CPU, CUDA GPU, air-gapped). Requirements: Python 3.11+, and a model backend (ollama, vLLM, MLX, or transformers).

Environments

SLM runs on four supported environments. The retrieval, grounding, UI, and evaluation layers are identical everywhere; only the model backend differs.

EnvironmentBackendBest forGPU
macOS (Apple Silicon)MLX or ollamaLocal single-user, LoRA trainingApple GPU (unified)
Linux (CPU)ollama / transformersSelf-hosted VM, no GPUNone needed
Linux + CUDAvLLM + PEFT/TRLProduction, multi-adapter, DPONVIDIA GPU
Air-gapped / offlineollama (local)Restricted networksOptional

Agents

Every pack is exposed as an OpenAI-compatible endpoint (/v1/chat/completions, /v1/models, /health). Any agent that speaks OpenAI's API can connect:

  • LangChain / LlamaIndex / AutoGen chat models
  • Custom copilots and internal tools
  • Slack / Discord / webhook bridges
  • CLI: curl http://localhost:8081/v1/chat/completions

Supported LLMs

The engine is model-agnostic. Verified backends and model families:

BackendModel familiesNotes
ollamaLlama 3.x, Qwen 2.5/3, Mistral, Gemma, DeepSeekSimplest local server; pull any GGUF
vLLMAny HuggingFace modelOpenAI-compatible; LoRA multi-adapter
MLXLlama, Qwen, Mistral (MLX format)Apple Silicon; LoRA training built in
transformers+PEFTAny HF CausalLMCPU serving; adapter loading
Judgeqwen2.5:3b (local) or any APIUsed for evaluation only

Serving

# OpenAI-compatible server for a pack (ollama backend)
$ python -m engine.serve.ollama my-corpus --port 8081 --model llama3.1:8b

# MLX backend (Apple)
$ make serve PACK=my-corpus PORT=8081

# Health check
$ curl localhost:8081/health
{"status":"ok","model":"my-corpus","loaded":true}

Evaluation

Every pack gets a closed-book evaluation on held-out sections, graded for correctness, hallucination rate, correct abstention, and latency. Results are written to packs/<name>/reports/ and surfaced as a quality grade in the UI.

$ make eval PACK=my-corpus
| variant | accuracy | halluc. | abstention | latency |
| base    |   0.00   |  0.13   |    -       |  0.5s   |
| trained |   1.00   |  0.00   |   0.75     |  1.6s   |

Runbook

  • Start: make ui (UI), make serve PACK=x PORT=8081 (model)
  • Add a pack: create folder + pack.yaml, put docs in raw/, make train PACK=x
  • Update a pack: add files, make setup (re-ingest), make datagen, make train
  • Roll back an adapter: point adapter_registry.json at the previous lora_N, restart the server
  • Back up: archive packs/*/{pack.yaml,adapters,data,processed} daily
  • Monitor: /admin/<pack> shows logs, eval tables, adapter versions

Known limitations

  • Longer books: retrieval is best-effort; borderline questions may surface a partial passage.
  • Word/Excel/PPT binary extraction is currently text-based (roadmap: full parsing).
  • The local 3B judge is the weakest link in evaluation; use a stronger judge for high-stakes grading.
  • Multi-adapter hot-swap requires vLLM (one adapter per process on MLX/ollama).
  • Weights can memorize training text — apply access control to adapters, not just the UI.
Something not covered? Contact sohil@qaso.ai.