What does model validation involve in practice?
Held-out data first: a model is judged on an eval dataset it never saw, because training-set performance flatters every model. Beyond accuracy, validation checks calibration, stability across segments, and behavior under drift once real data starts moving. Validation is a schedule, not an event: models degrade as the world changes, so the tests rerun for as long as the model runs, ideally inside an AI evaluation harness that makes rerunning cheap.
What are the main model validation techniques?
Splits, resampling and backtests, chosen by how the model will be used. A holdout split reserves data the model never trains on; cross-validation rotates that reserve so a small dataset still yields a stable estimate; backtesting replays history in time order, which is the honest option whenever the past predicts the future, as in credit or demand models. Beyond the headline accuracy number, the discipline checks performance per segment, since a model can be right on average and wrong for exactly one customer group, and it watches stability as fresh data drifts away from the training distribution.
Why do banks and insurers treat model validation as a control?
Because supervisors do. Model risk guidance in banking made independent validation a formal function with owners and review cycles, and insurers now hear the same from rating agencies: weak validation reads as weak governance. The catch of the moment is scope, since much of that guidance predates generative AI, leaving firms to extend the discipline themselves. Validating the model is still not validating the system: runtime checks like LLM output validation guard each answer, and AI assurance tests whether the whole control set supports the claims made for it.
What changes when the model is generative?
The ground truth disappears. A classifier is right or wrong; a drafted email or an agent’s plan is better or worse, so validation shifts from accuracy on a test set to scored behavior on a maintained suite of real tasks, often graded by LLM-as-a-judge with humans auditing the judge. Two habits carry over from the classic discipline and matter more here. Validate on data the model could not have memorized, the generative cousin of the holdout rule, because benchmark contamination flatters models the way training-set accuracy always did. And revalidate on every change, including the provider’s: a vendor’s silent model update is a new model in production, whether or not the version string moved.