Fine-tuned small models

Small models that know your domain. Weights included.

Where it beats retrieval alone, we fine-tune a small open-weight model on your domain, on your infrastructure. It runs on modest hardware, responds fast, and belongs to you.

Trained on
Your infrastructure
Owned by
You, weights included
Hardware
Modest compared to large models
Chosen when
It beats retrieval alone

In short

A fine-tuned small language model (SLM) is a compact open-weight model further trained on your own documents and tasks. It can match larger general models on narrow domain work while needing far less hardware. DeepVox fine-tunes on your infrastructure, and the resulting weights belong to you.

Why small

Four reasons to fine-tune.

Less hardware

Small models run on far fewer GPUs, which keeps on-prem deployments affordable.

Faster replies

Lower latency, which matters most for voice agents.

Your vocabulary

Your terminology, formats and house style, learned rather than prompted.

Yours to keep

The weights are delivered to you and run anywhere you choose.

Retrieval or fine-tuning

We pick on evidence, not fashion.

Most systems use retrieval; fine-tuning is added where it measurably helps. Often the answer is both.

SituationBetter approach
Facts change often (policies, prices, cases)Retrieval: answers stay current without retraining
Answers must cite a sourceRetrieval: the source is attached to every answer
A narrow, repeated task (classify, extract, format)Fine-tuned small model
Strict latency budget (voice)Fine-tuned small model, often with retrieval
Specialist vocabulary or house styleFine-tuned small model with retrieval

How we fine-tune

Measured before and after.

  1. Define the task

    In discovery we agree the task and an evaluation set built from your own examples.

  2. Benchmark first

    Retrieval alone and candidate base models are measured on that set.

  3. Fine-tune on your hardware

    Training runs in your environment; your data never leaves it.

  4. Prove it and hand over

    The fine-tuned model must beat the benchmark; then the weights and the evaluation set are yours.

FAQ

Questions, answered.

Is our data used to train anyone else’s models?

No. Fine-tuning happens on your infrastructure and the resulting model belongs to you. Nothing is sent to us or a model provider.

Which base models do you use?

Open-weight models such as Qwen and Mistral families, chosen in the mapping step against your task and hardware.

Does fine-tuning replace retrieval?

Rarely. Facts that change belong in retrieval; fine-tuning teaches tasks, formats and vocabulary.