Two ways to align an LLM to your needs, at very different data and cost complexity.
Start with SFT — it's cheaper, faster and gets you 80% of the way. Add DPO when preference data exists and the last 20% of quality matters.
Supervised fine-tuning trains the model on curated (prompt → ideal response) pairs. Simple, effective, well-tooled.
RLHF / DPO trains the model on preferences (response A preferred over response B). Better at nuanced tone, safety and calibration but needs preference data.
For most enterprise use cases, SFT alone is enough. Preference-based methods pay off for consumer-facing product quality and safety-critical alignment.
Train the model on labelled (input → output) pairs to internalise a task, format, tone or domain.
Train the model on preferences between responses, aligning it to what humans (or a reward model) prefer.
| Criterion | Supervised Fine-tuning (SFT) | Reinforcement Fine-tuning (RLHF / DPO) |
|---|---|---|
| Data needed | 500–10k input/output pairs | 1k–50k preference pairs |
| Cost / complexity | Low | Medium (DPO) – High (RLHF) |
| Best for | Task, format, domain | Tone, safety, helpfulness |
| Tool maturity | Very mature (TRL, Axolotl) | DPO mature; RLHF harder |
| Risk of quality collapse | Low | Real if over-trained |
| Time to production | Days | Weeks |
The typical enterprise arc: SFT to nail format and domain, then DPO on top with a few thousand preference pairs to lift subjective quality. Combining beats either alone.
Always first. Fine-tuning is right when prompting + retrieval have hit their ceiling — not before.
For most enterprises, no — DPO now delivers most of the benefit at a fraction of the complexity. RLHF is mostly justified at frontier-lab scale.
Yes — OpenAI, Anthropic and Google all support SFT (and some form of preference tuning). Open-source stacks give you full RLHF/DPO freedom.
Insights, use cases and industries that put this decision into context.
Every failed AI initiative we've audited failed on data — not models. The five data foundations we insist on before scoping a single copilot.
AI engineering is the discipline of turning models, data and tools into reliable business systems. Here's what it actually covers, how it differs from traditional software engineering, and where the ROI shows up.
pgvector, Pinecone, Weaviate, Qdrant, Milvus — a practical decision framework based on scale, latency, hybrid search, and total cost of ownership for enterprise RAG.
Deflect 60%+ of tier-1 tickets without hurting CSAT.
Straight-through processing for accounts payable — from PDF to ERP.
AI for firms that sell expertise
From AI prototype to production
Talk to a senior AI consultant from T7 about your industry, workflow, or product idea. Free, no commitment — reply within one business day.