Architecture · All markets

RAG or a fine-tuned small model? How we decide in discovery

In short

Retrieval-augmented generation (RAG) suits facts that change and answers that must cite a source. A fine-tuned small language model suits narrow, repeated tasks, specialist vocabulary and tight latency budgets such as voice. We decide by measuring both on an evaluation set built from the client's own questions, and often use both together.

Two different jobs

Retrieval finds the right passages in your documents at question time and gives them to a model to answer from. Answers stay current without retraining and can cite their source.

Fine-tuning further trains a model on your examples. The model learns tasks, formats and vocabulary. A fine-tuned small model can match much larger general models on narrow work while running on modest hardware.

Rules of thumb

  • Facts that change, such as policies, prices or cases: retrieval.
  • Answers that must cite a source: retrieval.
  • Repeated narrow tasks, such as classify, extract or format: fine-tuned small model.
  • Strict latency, such as voice: fine-tuned small model, often with retrieval.
  • Specialist vocabulary or house style: both.

How we decide

  1. Build an evaluation set from your own questions and tasks, with expected answers.
  2. Measure retrieval alone. It is the baseline and often enough.
  3. Try a fine-tuned small model where the task is narrow, repeated or latency-bound.
  4. Keep what wins, and keep the evaluation set to re-test every future change.

Ownership

Fine-tuning runs on your infrastructure, and the resulting weights belong to you. Keep personal data out of training sets unless there is a clear basis for it.

See also fine-tuned small models and RAG vs fine-tuning.

General information, not legal advice. Questions? Write to sales@deepvox.ai.