Fine-tuned small models
Small models that know your domain. Weights included.
Where it beats retrieval alone, we fine-tune a small open-weight model on your domain, on your infrastructure. It runs on modest hardware, responds fast, and belongs to you.
- Trained on
- Your infrastructure
- Owned by
- You, weights included
- Hardware
- Modest compared to large models
- Chosen when
- It beats retrieval alone
In short
A fine-tuned small language model (SLM) is a compact open-weight model further trained on your own documents and tasks. It can match larger general models on narrow domain work while needing far less hardware. DeepVox fine-tunes on your infrastructure, and the resulting weights belong to you.
Why small
Four reasons to fine-tune.
Less hardware
Small models run on far fewer GPUs, which keeps on-prem deployments affordable.
Faster replies
Lower latency, which matters most for voice agents.
Your vocabulary
Your terminology, formats and house style, learned rather than prompted.
Yours to keep
The weights are delivered to you and run anywhere you choose.
Retrieval or fine-tuning
We pick on evidence, not fashion.
Most systems use retrieval; fine-tuning is added where it measurably helps. Often the answer is both.
| Situation | Better approach |
|---|---|
| Facts change often (policies, prices, cases) | Retrieval: answers stay current without retraining |
| Answers must cite a source | Retrieval: the source is attached to every answer |
| A narrow, repeated task (classify, extract, format) | Fine-tuned small model |
| Strict latency budget (voice) | Fine-tuned small model, often with retrieval |
| Specialist vocabulary or house style | Fine-tuned small model with retrieval |
How we fine-tune
Measured before and after.
Define the task
In discovery we agree the task and an evaluation set built from your own examples.
Benchmark first
Retrieval alone and candidate base models are measured on that set.
Fine-tune on your hardware
Training runs in your environment; your data never leaves it.
Prove it and hand over
The fine-tuned model must beat the benchmark; then the weights and the evaluation set are yours.
FAQ
Questions, answered.
Is our data used to train anyone else’s models?
No. Fine-tuning happens on your infrastructure and the resulting model belongs to you. Nothing is sent to us or a model provider.
Which base models do you use?
Open-weight models such as Qwen and Mistral families, chosen in the mapping step against your task and hardware.
Does fine-tuning replace retrieval?
Rarely. Facts that change belong in retrieval; fine-tuning teaches tasks, formats and vocabulary.
Related
Knowledge & retrieval
The core most systems start with.
Learn more →RAG vs fine-tuning
A side-by-side guide.
Learn more →Deployment
Hardware sizing and hosting options.
Learn more →Start with two weeks of evidence, not a sales call.
A fixed-price discovery sprint, credited against whatever comes next. Or write to sales@deepvox.ai.