Skip to content

Fine-tuning use cases — a validation record (real 3B runs)

This page registers a real, GPU-backed validation of the fine-tuning path on the default qwen2.5-3b-instruct base — the model an operator actually ships. Every result and screenshot below comes from training real QLoRA adapters on the worker, scoring them on held-out data through the eval gate, and serving them from the console. It is reproduced by the opt-in suite tests/integration/test_lora_use_cases.py.

The matrix — use cases trained and results

Use case What it learns Gate result (3B) Served on a held-out input
Structured extraction text → a fixed JSON schema ✅ pass — score 0.33, +0.20 over base {"vendor": "Qorvex", "amount": 7700}
Classification / routing message → one label ✅ pass — score 0.67 (absolute bar) technical
Fixed format / voice reply in the house template ✅ pass — score 0.25, +0.22 over base "Sure thing! … Let us know if you need anything else."
Knowledge + behavior RAG facts + brand voice, one endpoint ✅ pass — +0.34 over base cites the doc (LunaBeans2026) and signs off in the trained voice
Garbage (negative control) nothing — patternless data blocked — delta −0.002 (cannot serve — gate refused it)

The negative control matters as much as the passes: a dataset with no learnable pattern must not clear the gate, and it doesn't. (An earlier control accidentally used a fixed sentence template — which is a learnable pattern — and a capable 3B correctly learned it; the fix was to make the control genuinely patternless, not to weaken the gate.)

1 · Structured extraction — invoice → JSON

The textbook fine-tune: teach the model to emit a fixed schema. The gate passes, and the served endpoint returns clean JSON for an invoice it never trained on.

Eval gate PASSED for the extraction adapter

Playground: an invoice is turned into {"vendor": "Qorvex", "amount": 7700}

2 · Fixed format / house voice

Teach the assistant to answer in one consistent template. Note the model paraphrases the opener ("Sure thing!") but keeps the learned house phrasing and the exact closer — the shape is what it learned.

Eval gate PASSED for the format adapter

Playground: a question answered in the trained house style

3 · Classification / routing

A support-ticket triage fine-tune (message → billing / technical / account). Full walkthrough with screenshots: Fine-tuning, proven.

4 · Knowledge + behavior on one endpoint

The production pattern — RAG facts (cited) composed with a fine-tuned brand voice in a single call. Full walkthrough: Knowledge + behavior together.

Reproducing this — and a note on small GPUs

# Whole matrix on the 3B (opt-in; heavy — real QLoRA per scenario). Cap the serving cache
# to ONE resident model so back-to-back serving of distinct 3B fine-tunes can't pile up:
docker compose exec -e ADAPTA_RUN_LORA_USECASES=1 -e ADAPTA_MAX_LOADED_MODELS=1 app \
  python -m pytest -m "integration and slow" tests/integration/test_lora_use_cases.py -s

# Faster, lower-VRAM floor run on the 0.5B:
docker compose exec -e ADAPTA_RUN_LORA_USECASES=1 -e ADAPTA_USECASE_BASE=Qwen/Qwen2.5-0.5B-Instruct \
  -e ADAPTA_USECASE_EPOCHS=25 app python -m pytest -m "integration and slow" \
  tests/integration/test_lora_use_cases.py -s

Serving cache size matters when one host serves many large fine-tunes

All five scenarios pass on the 3B in a single back-to-back run — verified green (5 passed). The catch: the serving cache keeps up to ADAPTA_MAX_LOADED_MODELS (default 2) distinct models resident, and on a host that serves inference on CPU (the GPU is reserved for training), holding two large 3B fine-tunes at once and loading a third mid-suite made the server stop responding. Setting ADAPTA_MAX_LOADED_MODELS=1 (now a compose passthrough) keeps one model resident at a time and the whole suite completes cleanly — that's the recommended setting for a single box that both trains and serves several large adapters. Separately, the eval gate loads its base in 4-bit, so a 3B evaluation fits comfortably on an 8 GB card and runs fast.