"Should we fine-tune a model on our data?" is one of the first questions business teams ask when starting an AI project. The intuition is understandable: our data is special, so the model should learn it. In practice, fine-tuning is the right first step far less often than people expect, and the wrong choice wastes months. This guide explains what each customisation method actually changes, what the evidence says, and a decision process you can apply to your project.

Three levers, three different effects

Lever What it changes Best for Time to first result
Prompting (incl. few-shot) The instructions and examples the model sees at runtime Behaviour, format, tone, task definition Hours
RAG The knowledge available at runtime via retrieval Facts, documents, frequently changing data, citations Days to weeks
Fine-tuning The model's weights Consistent style or format at scale, narrow skills, smaller/cheaper models, latency Weeks

A useful mental model: prompting tells the model what to do, RAG gives it what to know, fine-tuning changes how it tends to behave. Most real problems are about the first two.

Start with prompting

Modern models are capable enough that a well-structured prompt with a handful of examples solves a large share of business tasks: classification, extraction, drafting, summarisation, translation. Prompting has huge advantages: changes take minutes, you can switch models freely, and behaviour is inspectable. Before considering anything else, build a solid prompt and an evaluation set; the method is in prompt engineering for developers and LLM evals.

Prompting reaches its limits when the task needs knowledge the model does not have, when prompts become huge and expensive at high volume, or when the model cannot reliably follow a complex format even with examples.

Add RAG for knowledge

If the model needs your company's facts — product specifications, policies, contracts, tickets, prices — retrieval is almost always the right approach. RAG was introduced as a way to combine a model's parametric knowledge with a retrievable non-parametric memory (Lewis et al., 2020), and it has properties fine-tuning cannot match:

  • Freshness. Update a document and the next answer reflects it. A fine-tuned model is frozen at training time.
  • Citations. Answers can link to sources, which builds trust and makes errors checkable.
  • Access control. Retrieval can respect user permissions; weights cannot "forget" a document for a particular user.
  • Deletion. Removing data (for example, under GDPR) means deleting it from the index, not retraining.

The evidence also favours retrieval for knowledge. A study comparing knowledge injection methods found that RAG consistently outperformed unsupervised fine-tuning, both for knowledge seen during training and for entirely new knowledge, and that models struggled to learn new facts through fine-tuning alone (Ovadia et al., 2023). Other work suggests that fine-tuning on new facts can increase hallucination, because the model learns to produce confident answers beyond its actual knowledge (Gekhman et al., 2024).

Building good RAG is its own discipline — chunking, embeddings, hybrid search, reranking. See RAG for business, embeddings and chunking and vector databases.

When fine-tuning is the right tool

Fine-tuning shines when you need to change behaviour consistently and cheaply at scale:

  • Narrow, high-volume tasks. Classifying millions of support tickets or extracting fields from one document type. A fine-tuned small model can match a large general model on that one task at a fraction of the cost and latency.
  • Strict output style or format that prompts and structured outputs cannot enforce reliably — a brand voice, a domain-specific notation.
  • Shorter prompts. If every request carries 3,000 tokens of instructions and examples, fine-tuning can bake them in.
  • Domain language where the base model genuinely struggles: specialised terminology, under-represented languages or dialects — though here, test strong multilingual base models first.
  • Distillation. Use a large model to generate high-quality outputs, then fine-tune a smaller model on them for production.
  • Tool use patterns in agents with a fixed toolset, when you need a small model to call tools reliably.

Fine-tuning is the wrong tool for teaching facts that change, adding citation ability, enforcing permissions or fixing problems that better prompts would fix.

Fine-tuning methods in brief

Full fine-tuning updates all weights. It is expensive for large models and mostly unnecessary.

LoRA (Low-Rank Adaptation) freezes the base model and trains small low-rank matrices injected into its layers. For GPT-3 175B, LoRA reduced trainable parameters by 10,000x and GPU memory by 3x with comparable quality (Hu et al., 2021). Adapters are small files you can swap per task.

QLoRA combines LoRA with a 4-bit quantised base model, making it possible to fine-tune a 65B model on a single 48 GB GPU (Dettmers et al., 2023). This is what makes fine-tuning open-weight models practical on modest hardware.

Preference tuning (RLHF, DPO) teaches the model which of two outputs is better rather than imitating a single target. Useful for tone and style alignment, but needs preference data (Rafailov et al., 2023).

Hosted fine-tuning from model providers (for example OpenAI fine-tuning) handles infrastructure for you, with less control. For open-weight models, the Hugging Face PEFT library and tools built on it are the standard path.

Data: the real cost of fine-tuning

The compute is usually the cheap part. The expensive part is data:

  • Quantity: hundreds of examples can shift style; thousands are typical for reliable task performance; more for complex tasks.
  • Quality: each example must be what you want the model to produce. Inconsistent labels teach inconsistency. Reviewing 2,000 examples is weeks of expert time.
  • Coverage: examples must cover edge cases, refusals and the real distribution of inputs.
  • Hold-out evaluation: a separate test set you never train on, plus your existing eval suite.
  • Maintenance: when requirements change, you re-curate and re-train.
{"messages": [
  {"role": "system", "content": "Classify the support ticket. Output one label."},
  {"role": "user", "content": "Не можу завантажити акт звірки, кнопка сіра"},
  {"role": "assistant", "content": "documents.export_bug"}
]}

If you cannot produce a few thousand examples of this quality, fine-tuning is probably premature.

Combining the levers

The methods are not mutually exclusive. Common production combinations:

  • Prompting + RAG: the default for knowledge assistants, support bots and internal search.
  • RAG + fine-tuned model: a model fine-tuned to answer from retrieved context in the right format and to say "not in the documents" — the idea behind retrieval-aware fine-tuning such as RAFT (Zhang et al., 2024).
  • Large model → distilled small model: prototype with a frontier model and prompting, collect outputs, then fine-tune a small model for cost and latency, often self-hosted.
  • Routing: a fine-tuned small model handles the common cases; uncertain ones go to a large model. See LLM cost optimization.

A decision process

  1. Define success and build an eval set of 100+ real cases.
  2. Prompt a strong model. Measure.
  3. If failures are about missing knowledge → add RAG. Measure.
  4. If failures are about behaviour or format → improve prompts, examples and structured outputs. Measure.
  5. If quality is good but cost or latency is too high at your volume → consider fine-tuning a smaller model (distillation). Measure against the same eval.
  6. If quality plateaus on a narrow task despite good prompts → consider fine-tuning with curated data.

Each step should be justified by eval results, not intuition. Model releases also move the baseline: a fine-tune that beat last year's base model may lose to this year's with a good prompt — see choosing an LLM.

FAQ

Can we fine-tune a model on all our company documents so it "knows" them? Technically yes, but it is the least effective way to give a model knowledge: poor recall of specific facts, no citations, no updates without retraining. Use RAG.

Is fine-tuning required for Ukrainian? Usually not. Strong commercial and open multilingual models handle Ukrainian well. Fine-tuning may help for a specific domain style or for small models.

How much does fine-tuning cost? LoRA on an open 7–14B model can cost tens to hundreds of dollars in GPU time. Data preparation and evaluation typically cost far more in people's time.

Does fine-tuning help with hallucinations? It can teach a model to refuse or to stick to provided context, but fine-tuning on new facts can increase hallucination. Grounding with RAG and verification is more reliable.

Sources

  1. Lewis et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.
  2. Ovadia et al. (2023). Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs.
  3. Gekhman et al. (2024). Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
  4. Hu et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models.
  5. Dettmers et al. (2023). QLoRA: Efficient Finetuning of Quantized LLMs.
  6. Rafailov et al. (2023). Direct Preference Optimization; Zhang et al. (2024). RAFT: Adapting Language Model to Domain Specific RAG.
  7. OpenAI. Fine-tuning guide; Hugging Face. PEFT.