Do you actually need to fine-tune a model?
Probably not first. Most problems people reach for fine-tuning to solve, the model does not know our data, the model does not follow our format, are solved faster and more cheaply by better context engineering: retrieving the right facts through RAG, or simply describing the format in the prompt with a structured output schema. Fine-tuning is the expensive, slow-to-iterate option, best saved for what context alone cannot fix, a consistent voice across thousands of calls, or a narrow classification task run at a cost a frontier model cannot justify, which is the gap model distillation targets by training a smaller student model on a larger teacher’s outputs.
What does fine-tuning actually change?
It bakes a pattern into the model’s weights instead of supplying it fresh at call time. That trade cuts both ways: the pattern survives without a prompt reminding it every call, but it also stops adapting the moment your data or requirements shift, and retraining means a new dataset, a new training run, and a new evaluation pass, not a one-line prompt edit.