Deployment
Runs where your data already lives.
On your own servers or in your private cloud, operated by your team or by us. Everything is containerised, documented and handed over with runbooks.
- Where
- On-premise or private cloud
- Operated by
- You, or us if you prefer
- Packaging
- Kubernetes and Helm
- Handover
- Runbooks and training
In short
DeepVox deploys the whole AI system, including connectors, index, models, voice and logs, on infrastructure you control: your data centre or your own private cloud account. Your team can operate it with our runbooks, or we can run it under a managed support agreement.
Options
Three ways to run it.
On-premise
Your data centre and GPUs. Nothing leaves your network.
Private cloud
Your own cloud account and region, under your keys and policies.
Start hosted, move later
Begin with a hosted model through the gateway and move inference on-prem without rebuilding.
What gets deployed
Every layer, in your environment.
| Layer | Typical components |
|---|---|
| Ingestion | Docling, Unstructured, Apache Tika, Tesseract |
| Index and vector store | PostgreSQL with pgvector, Qdrant, OpenSearch |
| Model serving | vLLM, SGLang, Ollama |
| Model gateway | LiteLLM, with usage and cost per department |
| Identity | Keycloak, Open Policy Agent, Entra ID integration |
| Observability | Langfuse, Phoenix, Grafana stack |
| Orchestration | n8n, custom services, Kubernetes and Helm |
Operated your way
Run it yourself, or let us.
- Runbooks and documentation as contracted deliverables.
- Training for your operations team.
- Managed support with a 99.9% uptime SLA, if you want it.
- A rehearsed exit test so a successor supplier can take over.
FAQ
Questions, answered.
What hardware do we need?
It depends on the models, users and whether you run voice; we size it in discovery. Small fine-tuned models keep the footprint modest.
Can we start without buying GPUs?
Yes. Start with a hosted model through the gateway and move on-prem later; it is a configuration change.
Who patches and upgrades it?
Your team with our runbooks, or us under managed support.
Related
Security
Controls and data handling.
Learn more →Fine-tuned small models
Smaller models, smaller footprint.
Learn more →Pricing
Infrastructure at cost, not per seat.
Learn more →Start with two weeks of evidence, not a sales call.
A fixed-price discovery sprint, credited against whatever comes next. Or write to sales@deepvox.ai.