Tailr
← All posts

What Is Fine-Tuning? How It Works and When to Use It Over RAG

· updated

Three steps of fine-tuning: collect examples, train a base model, then evaluate the tuned model

Fine-tuning comes up in almost every conversation about building with AI, usually right next to RAG and prompt engineering. It’s often the wrong tool, and sometimes the only right one. This guide explains what fine-tuning is, how it works, what LoRA means, and how to decide whether you need it at all.

The short answer: Fine-tuning is training an already-trained model a little more on your own examples so it gets better at one specific task, style or format. The model’s weights change, so the new behaviour is built in. The usual order of things to try is:

  1. Prompting: clear instructions and a few examples in the prompt.
  2. RAG: fetch the right documents and add them to the prompt.
  3. Fine-tuning: only when the first two can’t get the behaviour, format or cost you need.

Fine-tuning teaches a model how to respond. RAG gives it what to know.

How fine-tuning works

A large language model is first pre-trained on huge amounts of text, which is where it learns language and general knowledge. That’s expensive and done by a handful of labs.

Fine-tuning starts from that finished model and keeps training it on a much smaller dataset of your own: usually pairs of inputs and the ideal outputs. Each example nudges the model’s weights so its answers look more like yours.

A concrete example

Say you want a model to turn messy support emails into a strict JSON ticket with category, priority and summary. Prompting gets it right most of the time, but it sometimes adds chatty text or invents categories. You collect 800 real emails with the correct ticket for each, fine-tune a small model on them, and it now outputs the exact format every time, faster and cheaper than the large model you were prompting.

What fine-tuning data looks like

Most fine-tuning today uses the same chat format the model already speaks: a list of messages, ending with the answer you want it to learn. Each example is one line in a JSONL file (one JSON object per line). For the support-ticket example, one line might look like this, spread out here so it’s readable:

{"messages": [
  {"role": "system", "content": "Turn support emails into a JSON ticket."},
  {"role": "user", "content": "Hi, I was charged twice for my March invoice and need a refund asap. Thanks, Sam"},
  {"role": "assistant", "content": "{\"category\": \"billing\", \"priority\": \"high\", \"summary\": \"Double charge on March invoice; refund requested\"}"}
]}

A few rules make the difference between a dataset that works and one that doesn’t:

  • Be consistent. If two similar emails get different categories, the model learns that the categories are random. Write a short labelling guide and stick to it.
  • Cover the edge cases. Include the awkward inputs: empty emails, three issues in one message, angry customers, other languages. The model only learns what it sees.
  • Match real inputs. Train on the kind of text the model will actually receive, not cleaned-up versions.
  • Keep a holdout set. Put 10 to 20% of examples aside and never train on them. That’s how you’ll know if it really learned.
  • Check for private data. Strip names, emails and account numbers you don’t need. Whatever goes into training can come back out.

A useful trick: write the first 50 examples by hand, use a strong model to draft the next few hundred in the same style, then review every one. Reviewing is much faster than writing, and it keeps the quality up.

The main types of fine-tuning

Type What it does When it’s used
Supervised fine-tuning (SFT) Trains on input-and-ideal-output pairs The most common kind: formats, tasks, tone
Preference tuning (RLHF, DPO) Trains on pairs of better and worse answers Making responses more helpful or on-brand
Full fine-tuning Updates every weight in the model Big budgets, big datasets
LoRA / QLoRA Trains a small add-on while the model stays frozen Most practical projects today

LoRA, explained simply

Full fine-tuning updates billions of numbers, which needs a lot of GPU memory. LoRA (low-rank adaptation) freezes the original model and trains a small set of extra weights, often well under 1% of the size, that sit alongside it. You get most of the benefit at a fraction of the cost, and you can keep several LoRA “adapters” for different tasks on top of one base model.

QLoRA loads the base model in 4-bit precision first, so even fairly large open models can be tuned on a single GPU.

Fine-tuning vs RAG vs prompting

Prompting RAG Fine-tuning
Changes the model? No No Yes
Best for Most tasks, quick iteration Fresh or private knowledge, citations Consistent format, style, narrow skills
Update knowledge Edit the prompt Update the documents Retrain
Upfront effort Minutes Days Days to weeks
Running cost Can be high with long prompts Moderate Often lower, with a smaller model

The most common mistake is fine-tuning to add facts. If the model doesn’t know your product docs, fine-tuning on them gives you a model that sounds confident and still gets details wrong, and it’s stale the moment the docs change. That’s a job for RAG.

Real-world fine-tuning examples

These are the kinds of jobs where teams actually reach for fine-tuning:

  • Classification at scale. Routing millions of support tickets, reviews or documents into categories. A small tuned model is fast and cheap enough to run on every item.
  • Structured extraction. Pulling the same fields out of invoices, contracts or medical notes into a fixed format, every time.
  • House style. A news site or brand that wants every summary or product description to follow its voice and length rules without a two-page style guide in each prompt.
  • Distillation. Using a large model to generate high-quality answers for a narrow task, then training a much smaller model on those answers so it can do the same job at a fraction of the cost and latency.
  • Domain language. Models for fields with heavy jargon, such as legal or clinical text, where a general model handles the vocabulary clumsily.
  • On-device and private models. Small open models tuned for one task so they can run on a phone, a laptop or a company’s own servers without sending data anywhere.

Notice what’s missing: “answer questions about our latest docs.” That’s knowledge, and it changes. RAG handles it better.

When fine-tuning is the right call

  • You need an exact output format every single time.
  • You want a consistent voice or style that’s hard to describe in a prompt.
  • You want to swap a big model for a small one that’s cheaper and faster on one narrow task.
  • Your prompt has become enormous with examples and rules, and you’re paying for it on every call.
  • You need the model to run privately on your own hardware.

How to fine-tune a model, step by step

  1. Define success first. Write 50 to 100 test cases and decide how you’ll score them. Without this you can’t tell if tuning helped.
  2. Get a baseline. Run your best prompt (and RAG, if relevant) on those test cases.
  3. Collect examples. Usually a few hundred to a few thousand clean input-and-output pairs. Consistency beats volume.
  4. Pick a route. A hosted fine-tuning API from a model provider, or open models with tools like Hugging Face libraries (huggingface.co) or Unsloth (unsloth.ai) for LoRA.
  5. Train, then evaluate on test cases the model never saw in training.
  6. Compare to the baseline. Keep the tuned model only if it clearly wins on quality, cost or speed.

How to evaluate a fine-tuned model

Evaluation is where most fine-tuning projects quietly fail. “It looks better” isn’t a result. Measure it:

  1. Use the holdout set. Run the tuned model and your baseline (best prompt, plus RAG if you use it) on the same examples the model never trained on.
  2. Score automatically where you can. For formats, check that the output parses and every field is valid. For classification, measure accuracy per category, not just overall, because a model can look great on average while failing a rare but important category.
  3. Use a rubric for open-ended output. For summaries or tone, write three to five criteria and score a sample by hand, or use a strong model as a judge and spot-check its scores.
  4. Check what you might have broken. Fine-tuning on a narrow task can make a model worse at things it used to do well. Test a handful of general prompts too.
  5. Compare cost and speed. A tuned small model that matches the big model’s quality at a fifth of the cost is a win, even if it isn’t more accurate.

Write the results down in a simple table: baseline vs tuned, quality, cost per 1,000 requests, and response time. That table is also exactly what an interviewer wants to hear about.

What fine-tuning costs

Costs have dropped a lot, but they come in three parts:

  • Data. Usually the biggest cost: the hours spent collecting, labelling and checking examples.
  • Training. With LoRA on a small open model, a run can cost a few dollars of rented GPU time. Hosted fine-tuning APIs charge by the amount of training text, and bigger models cost more.
  • Running it. Hosted tuned models are often priced a little higher per token than the base model; self-hosting means paying for a server. A smaller tuned model can still be far cheaper than prompting a large one with a long prompt.

Then there’s upkeep. When a better base model comes out, or your categories change, you’ll need to retrain. Budget for that from the start.

Common fine-tuning mistakes

Mistake Fix
Fine-tuning to teach facts Use RAG for knowledge
No evaluation set Write test cases before you train
Messy, inconsistent examples Clean and standardise them first
Testing on training data Hold out examples the model never sees
Forgetting upkeep Plan to retrain when the base model changes

Fine-tuning terms, quickly

  • Base model vs instruct model: a base model just continues text; an instruct (or chat) model has already been tuned to follow instructions. Most projects start from an instruct model.
  • Weights: the billions of numbers inside a model that fine-tuning adjusts.
  • Epoch: one full pass through your training data. A few epochs is typical; too many leads to overfitting.
  • Learning rate: how big each adjustment is. Too high and the model forgets what it knew; too low and it barely changes.
  • Overfitting: the model memorises the training examples instead of learning the pattern, so it does well on them and badly on new inputs.
  • Catastrophic forgetting: the model gets worse at general tasks after being tuned on a narrow one.
  • Adapter: the small set of extra weights LoRA trains, which can be swapped in and out.
  • Distillation: training a smaller model to copy a larger model’s outputs on a task.

Why fine-tuning matters for your career

If you’re heading into AI engineering, interviewers love asking “When would you fine-tune instead of using RAG?” A clear, honest answer, ideally backed by a small project where you measured the difference, says more than any certificate. Our list of AI engineer interview questions covers this and related topics.

On your resume, describe the result, not the technique alone: “Fine-tuned a small open model with LoRA to classify support tickets, matching the larger model’s accuracy at a fifth of the cost.” Tailr can help you line up projects like that with what a specific job listing asks for, tailoring your resume from the listing you’re viewing.

Try Tailr

Conclusion

Fine-tuning trains an existing model further on your own examples so it reliably does one thing your way. It’s great for formats, style and making small models punch above their weight, and poor at teaching facts. Start with prompting, add RAG for knowledge, and fine-tune only when you have a clear evaluation showing it’s worth it.

Frequently asked questions

01What is fine-tuning in simple terms?

Fine-tuning means taking a model that's already been trained and training it a little more on your own examples, so it gets better at a specific task, style or format. The model's weights change, so the new behaviour is built in rather than explained in every prompt.

02What is the difference between fine-tuning and RAG?

Fine-tuning changes how the model behaves by updating its weights with examples. RAG leaves the model as it is and adds relevant documents to the prompt at the moment you ask. Use fine-tuning for behaviour and format, and RAG for knowledge that changes or must be cited.

03How much data do you need to fine-tune a model?

It depends on the task, but useful results often start with a few hundred to a few thousand high-quality examples. Quality matters far more than quantity: a small set of clean, consistent examples beats a large set of messy ones.

04What is LoRA?

LoRA (low-rank adaptation) is a way to fine-tune a model by training a small set of extra weights while leaving the original model frozen. It needs far less memory and compute than full fine-tuning, and QLoRA goes further by loading the base model in 4-bit precision so it fits on a single GPU.

05Is fine-tuning expensive?

It's much cheaper than it used to be. With LoRA or a hosted fine-tuning API, a small model can be tuned for anywhere from a few dollars to a few hundred, depending on its size and the amount of data. The bigger cost is usually preparing good data and evaluating the results.

06When should you not fine-tune?

Don't fine-tune when better prompts, a few examples in the prompt, or RAG already solve the problem, or when the main issue is missing or frequently changing facts. Fine-tuning also adds upkeep, since you'll need to retrain when the base model or your requirements change.