Almost every client conversation about “customizing” an AI feature starts with the word “fine-tune.” Most of those conversations end somewhere else — usually a better prompt, sometimes a retrieval layer, and only occasionally an actual training run. That order matters, because it's also the order of cost and flexibility, from cheapest and easiest to change, to most expensive and hardest to undo.
Three tools that solve three different problems
It helps to be precise about what each option actually changes, because they aren't interchangeable:
Prompt engineering changes the instructions. Fastest to test, free to change, and the right first move for almost anything that looks like a formatting or instruction-following problem.
Retrieval (RAG) changes what the model knows, by pulling relevant documents in at request time. This is the fix for missing or changing knowledge — and it keeps facts current and citable, which fine-tuning never does.
Fine-tuning changes how the model behaves, by adjusting its weights on examples of the behavior you want. It's the right tool only for entrenched style, format, or domain habits that prompting genuinely can't hold.
The single most common mistake we see is reaching for the third option to solve the second problem — fine-tuning a model on internal documentation to “teach it the facts.” That bakes today's knowledge into frozen weights. It goes stale the moment anything changes, and unlike a retrieval index, you can't patch it by updating a document.
When fine-tuning actually earns its cost
Fine-tuning has a real, recurring bill: hosting the tuned model (or a premium on a hosted platform), on top of the one-time training run. That's only worth paying when:
The task is high-volume and stable — the same narrow job, run enough times that a small per-request improvement adds up to real savings.
Prompting has genuinely plateaued — you've tried and it still won't reliably hold the format, tone, or behavior you need.
The win is measurable — fewer tokens per request, a meaningfully higher success rate, or a capability you can't reach any other way.
If your requirements are still shifting, or your volume is low, the training cost never really amortizes, and every requirement change means retraining. That's a sign to stay with prompting or retrieval.
If you do fine-tune, LoRA is almost always the right shape
Full fine-tuning updates every weight in the model, which is powerful and, for anything at frontier scale, impractical for most teams. Parameter-efficient methods like LoRA freeze the base model and train small adapter matrices instead — a fraction of a percent of the parameters — for most of the benefit at a fraction of the compute and memory. That's why the large majority of practical fine-tunes today are LoRA rather than full retrains.
Open-weights models have also changed the economics here. A model like Thinking Machines' Inkling ships specifically designed to be fine-tuned on infrastructure you control, which means you own the resulting artifact instead of renting a tuned result inside a vendor's platform on the vendor's terms.
Where fine-tuning projects actually go wrong
Overfitting: too many epochs or too high a rank, and the model memorizes the training set instead of generalizing. Watch validation loss, not just training loss, and stop early.
No real evaluation: judging results by feel, or by the same examples used in training. Both are misleading. A held-out test set with a fixed metric is the only honest read.
Skipping the baseline: if you never measured how far prompting and retrieval alone got you, you can't actually prove the fine-tune bought anything.
A messy dataset: quality beats volume by a wide margin — a few hundred clean, consistent examples outperform tens of thousands of noisy ones. Read a sample of your own training data before you trust it.
The takeaway
Fine-tuning isn't wrong, it's just usually premature. Work down the ladder — prompt, then retrieval, then fine-tuning — and stop at the first rung that actually solves the problem. When we scope AI features for clients, proving that the cheaper options fall short is part of the deliverable, not a step we skip on the way to a training run.
