Fine-tuning: retraining an AI model on your own examples
The first question about fine-tuning is not how, but whether. Fine-tuning means taking a model that has already been trained and training it further on your own examples, so that it performs a specific task, format or tone more reliably. Google’s glossary calls it a second, task-specific training pass on a pre-trained model.
Three ways to change what a model does
| Approach | What changes | Suited to |
|---|---|---|
| Prompting with examples | Nothing in the model, only the instructions. | trying out a task, small volumes |
| Retrieval-augmented generation, based on embeddings | Nothing in the model; relevant documents are supplied with each request. | knowledge that changes, answers from your own documents |
| Fine-tuning | The model’s weights. | one job repeated at high volume |
OpenAI names three advantages of fine-tuning over prompting alone: you can provide more examples than fit into a single request, shorter prompts reduce cost and latency at scale, and you can train on proprietary data without sending it along with every request.
The main methods
- Supervised fine-tuning: examples of prompts paired with the correct response.
- Direct Preference Optimization (DPO): for each prompt, a better and a worse response.
- Reinforcement fine-tuning: the model answers, the answers are graded, and higher-scoring reasoning is reinforced.
- Vision fine-tuning: the same idea applied to image inputs.
These are the methods OpenAI documents for its own models. For open models, LoRA, short for low-rank adaptation, has become a common way to keep costs down: it freezes the original weights and trains small additional matrices inserted into the layers, an approach published by Hu and colleagues in 2021.
A scenario that goes wrong
Suppose a company fine-tunes a model on three years of its own support replies, so that answers sound like the team. Some of those replies quote prices that have since changed, and a few contradict each other. The tuned model repeats both, fluently, because nothing in the training told it which replies were wrong.
Nobody notices at first, because nobody held back a set of test questions to compare the tuned model with the original. A year later a new generation of the base model arrives, and the tuning has to be repeated on it. The replies also contained customer names, so the training data needed a legal basis under data protection law before any of this began. And if the company later offers the tuned model to others, the company’s role under the EU AI Act may change, which is worth checking before release.
A cheap first step
Before deciding anything, write down twenty real inputs and the outputs you expect for each. That list is the test set for prompting, retrieval and fine-tuning alike, and it often shows that a well-built prompt already gets most of the way. Working out which approach fits, including whether a local model makes sense, is part of an AI implementation.
Related terms
| Term | What it means |
|---|---|
| Base model | The pre-trained model that fine-tuning starts from. |
| Evaluation set | Examples held back from training to judge whether a change actually helped. |
| LoRA and QLoRA | Resource-efficient fine-tuning methods. |
Sources
- Google for Developers, Machine Learning Glossary, entry “fine-tuning”, accessed 14 September 2026
- OpenAI, Model optimization, accessed 14 September 2026
- Hu et al., “LoRA: Low-Rank Adaptation of Large Language Models”, arXiv, 2021
This page expands an entry from the minoka AI glossary, which covers many more terms in brief.
Frequently Asked Questions
Frequently Asked Questions
When is fine-tuning worth it?
When one narrowly defined job recurs at high volume, its output has to stay consistent, and prompting with examples either does not reach the required reliability or becomes too expensive.
How much data do I need for fine-tuning?
There is no fixed number. Consistent, correct examples matter more than volume, and you need a separate set of test examples to tell whether the tuned model is better than the original.
Can fine-tuning keep a model’s knowledge current?
Not well. Every change to the facts would mean another training run, so current information is better supplied at request time through retrieval-augmented generation.
What does Direct Preference Optimization do differently?
Instead of one correct answer per prompt, it trains on pairs of a better and a worse response, so the model learns which of two answers is preferred.
Preferred source
Prefer minoka.de on Google
If you add minoka.de as a preferred source, Google shows articles from this site more often in Top Stories, AI Overviews and AI Mode. One click, a Google account, revocable at any time in your source settings.
Prefer on GoogleOpens the source settings at Google
