โถ Want to run it yourself? Download the Python notebook โ it
recomputes every number on this page, on the same 400 real spectra (Windows-safe, ~2 min):
OSSL_Validation_Methods_WIN.zip
The one idea: a validation method is not a formula for accuracy โ it is a
protocol for hiding data. Each of the five hides data differently, so each answers a
different question. Same data, same model, different hiding rule โ different number.
Setup โ what we're working with
400 soil samples from the Open Soil Spectral Library (OSSL) v1.2. Each sample is a
reflectance spectrum, and we predict soil organic carbon (SOC) from it.
samples 400
bands 1051 (400โ2500 nm)
target log10(SOC %)
SOC raw range 0.05โ62.37 %
model PLS, 10 components
Drawn from 4 real libraries:
LUCAS.SSL 242, KSSL.SSL 133, ICRAF.ISRIC 23, LUCAS.WOODWELL.SSL 2.
How to read this: left โ 12 of the 400 real spectra; each line is one
soil's light fingerprint. Right โ SOC on a log scale, because raw SOC is heavily right-skewed
(a few soils are very carbon-rich) and that skew wrecks regression.
All five on one page
How to read this: blue = training data, orange = test,
green = validation. Every method below is just a different way of painting this picture.
How to read this: the first seven bars are all in the same range โ
they are all "random" hiding rules. The last red bar is a fundamentally different question.
Keep that gap in mind; section 7 explains it.
1 ยท Holdout Validation
Rule: split once (here 70/30), train on the big part, test on the small part, report
that one number.
The trap: the answer depends on which samples happened to land in
the test set. Five different random splits, identical model, identical data โ Rยฒ from
0.694 to 0.796.
How to read this: left โ five single splits; the bar heights are
the verdicts from the same model. Right โ 20 splits, and the mean sits in the middle with a
real spread around it.
1
You get an unbiased estimate โ nothing about
the test rows leaked into training.
2
But with high variance: the number is a
lottery draw. Report one split and you can't tell if the model is good or you got lucky.
Analogy: it's like judging a student on one randomly chosen exam
question. The question is fair, but the verdict swings wildly depending on which question they
got.
2 ยท Repeated Holdout Validation
Rule: do the same thing many times with different splits, then report the
mean and the spread. Cheapest possible fix for Method 1's variance.
On our data, 20 repeats give Rยฒ = 0.723 ยฑ 0.036
(range 0.656โ0.786). Note the mean barely moved versus a single split โ
what you gained is the error bar.
Analogy: instead of one exam question, you give the student 20
randomly drawn questions and average. The average is a fairer verdict, and the spread tells you
how much the verdict depends on the draw.
3 ยท Bootstrap Validation
Rule: draw n rows with replacement โ some rows appear several
times, and about 37% never appear at all. Train on the drawn rows, then score the ones
you missed. Those leftovers are the out-of-bag (OOB) samples.
How to read this: left โ one resample of 400 rows; 143 rows were
never drawn (those are your free test set). Middle โ the optimism problem: scoring a
model on its own training rows (apparent error 0.2942) looks better than
reality (OOB 0.3208). Right โ the spread across all
200 resamples.
The .632 estimator blends the two:
0.368รapparent + 0.632รOOB = 0.3110. The exact split of the weights
comes from the bootstrap theory of optimism โ you don't need to memorise it, just know it exists
to correct the too-rosy apparent error.
The subtlety: because rows are duplicated in the in-bag set, the model
partly memorises them. That is exactly why you must score on the OOB rows and never on the
training rows.
4 ยท 3-Way Holdout
Rule: three parts, three jobs โ train fits the model, validation picks
the hyperparameters, test is touched once at the very end to report.
Here we tune one thing: how many PLS components to use (grid 1โ20).
How to read this: left โ error curves for all three sets; the
validation curve picks ncomp = 15. Right โ if you instead let the test set
pick its own best model (ncomp = 7), the reported error drops to
0.3230 instead of the honest 0.3591.
That 0.0361 RMSE gap is the contamination. It is not a bug โ it is what
happens when the test set stops being a neutral examiner and starts being part of the design.
Analogy: practice exams (validation) vs the final exam (test). If the
teacher writes the final after seeing which questions you got right, your final score
flatters you.
5 ยท Cross-Validation (k-fold, and Leave-One-Out)
Rule: split into k folds; train on kโ1, test on the one left
out; rotate so every row is tested exactly once; average the k scores.
How to read this: left โ per-fold scores; the box height is the
disagreement between folds. Right โ colour shows which fold tested each row, after sorting by
SOC; every row is covered exactly once.
| Scheme | Rยฒ (mean) | sd across folds | RMSE (mean) |
| 5-fold | 0.725 |
0.052 | 0.3229 |
| 10-fold | 0.722 |
0.087 | 0.3196 |
| Leave-one-out (k = n = 400) | 0.737 |
โ | 0.3203 |
Counter-intuitive bit: more folds is not automatically better.
LOO trains on nโ1 rows every time, so its n folds are nearly identical โ the mean is good but
the variance of the estimate can be understated. It also costs 400 model fits instead
of 5. Note 10-fold here has a wider spread (0.087) than 5-fold
(0.052) โ smaller test folds means noisier individual scores.
How to read this: left โ predicted vs observed; right โ residuals.
The far-right panel is the useful one: predictions go wrong at very low and very high SOC,
which is a data-coverage problem, not a validation-method problem.
6 ยท All five side by side
| Method | Rยฒ | RMSE | Note |
| Holdout (single split) | 0.744 | 0.3315 | one number, high variance |
| Repeated holdout ร20 | 0.723 | โ | sd Rยฒ 0.036 |
| Bootstrap OOB | 0.737 | 0.3208 | 200 resamples |
| 3-way (test, honest) | 0.704 | 0.3591 | ncomp=15 from validation |
| 5-fold CV | 0.725 | 0.3229 | sd Rยฒ 0.052 |
| 10-fold CV | 0.722 | 0.3196 | sd Rยฒ 0.087 |
| Leave-one-out | 0.737 | 0.3203 | 400 fits |
| Leave-one-DATASET-out | 0.203 | 0.4551 | the honest one |
7 ยท The twist that matters for soil data
Every method above hides rows at random. But two soil samples from the same region are
usually similar, so a random split can leave near-copies of your test samples inside the
training set. Your score then partly measures memorisation of neighbours, not the ability
to predict a new place.
The honest version holds out an entire source library: train on the others, predict a
library the model has never seen.
How to read this: left โ holding out a whole library gives
0.762 (KSSL), 0.064 (ICRAF), and
-0.215 (LUCAS). The dashed line is what random 5-fold claimed.
Right โ the gap: 0.725 vs 0.203.
Read this carefully: Rยฒ = 0 means "no better than always predicting the
average". A negative Rยฒ means worse than that. LUCAS scored
-0.215 โ the model trained on American and African soils actively
mispredicts European soils. Random 5-fold hid that completely behind a score of
0.725.
The lesson: the leakage gap is
0.522 Rยฒ. The hiding rule must match the question you
actually care about. "A new sample from the same places" and "a new place, new instrument,
new lab" are very different claims โ and only one of them survives this test.
๐ Quick answer sheet
| Method | How it hides data | Answers the question | Watch out for |
| Holdout | one random split | "a new sample, same population" |
high variance โ one lucky draw |
| Repeated holdout | many random splits | same, with an error bar |
cost scales with repeats |
| Bootstrap | resample with replacement; score OOB |
same, reusing data efficiently | apparent error is way too optimistic |
| 3-way holdout | separate set for tuning | same, with honest tuning |
tuning on test contaminates the score |
| k-fold CV | rotate k test folds | same, every row tested once |
fold spread; smaller folds are noisier |
| LOO (k = n) | leave exactly one row out | same, maximum training data |
n fits; folds nearly identical โ variance understated |
๐ง Checklist before you report a score
1
Is it one split or a distribution? If one,
report the spread too.
2
Was the test set ever used for decisions
(tuning, feature selection, early stopping)? If yes, the score is contaminated.
3
Are you scoring on training rows anywhere?
Remember the apparent-vs-OOB gap of
0.0266 RMSE.
4
Do your folds respect the structure of the
data (region, site, instrument, year)? If not, you are measuring interpolation, not
generalisation.
5
Does your validation rule match the claim you're
making in the paper?