Two paths to production LLM. The right answer depends on volume, sensitivity and team capacity.
Managed API for most enterprises under 100M tokens/day. Self-hosted for extreme volume, regulated air-gap, or genuine model customisation.
Managed APIs (OpenAI, Anthropic, Google) win on accuracy at frontier tasks, ops simplicity and speed to production. You trade unit economics at very high volume.
Self-hosted open-source (Llama 3.3, Mistral, Qwen) wins on unit cost at scale, full data control and deep customisation. You take on GPU ops, model updates and eval.
In practice, most enterprises are still better served by managed APIs. Self-hosting pays off at extreme volume or when regulation genuinely rules out managed.
Open-source models running on your GPUs — full control, custom fine-tuning, no per-token vendor bill.
Frontier models exposed as APIs with enterprise tenancy options.
| Criterion | Self-hosted (Llama / Mistral) | Managed API (OpenAI / Anthropic) |
|---|---|---|
| Accuracy (hard tasks) | Very strong (behind frontier) | Best-in-class |
| Cost at 10B tokens/mo | Low (amortised GPUs) | Very high |
| Ops burden | High (GPU + MLOps) | Near zero |
| Data control | Full (air-gap possible) | Zero-retention available but data leaves |
| Customisation | Full fine-tune / LoRA | Limited (prompt, RAG, some fine-tune) |
| Time to production | Weeks–months | Days |
The dominant pattern in 2026: managed APIs for reasoning-heavy queries, self-hosted small models for high-volume classification, extraction and routing. Model routing gets 70–90% of self-hosting's cost win at 5% of the ops burden.
Yes — OpenAI, Anthropic and Google all offer fine-tuning. It's more limited than self-hosted LoRA but often good enough.
GPU rent for a Llama 3.3 70B endpoint at moderate load: $5–15k/month. Add MLOps, retraining and updates — real all-in is often 2–3x list.
That's the recommended default in 2026. Route by query type. Best economics with minimum ops.
Insights, use cases and industries that put this decision into context.
Model routing, prompt compression, caching, distillation and eval-driven downgrades — the levers we use to bring enterprise LLM bills under control without hurting quality.
How we cut a customer's monthly LLM bill 84% without touching accuracy — routing, distillation, caching and the boring engineering behind every dollar.
pgvector, Pinecone, Weaviate, Qdrant, Milvus — a practical decision framework based on scale, latency, hybrid search, and total cost of ownership for enterprise RAG.
Deflect 60%+ of tier-1 tickets without hurting CSAT.
A copilot that drafts, researches and updates the CRM — so reps sell.
From AI prototype to production
AI-powered learning platforms
Talk to a senior AI consultant from T7 about your industry, workflow, or product idea. Free, no commitment — reply within one business day.