๐Ÿงช The Five Validation Methods

All five, run on the same 400 real OSSL soil spectra โ€” so you can see why they disagree

โ–ถ Want to run it yourself? Download the Python notebook โ€” it recomputes every number on this page, on the same 400 real spectra (Windows-safe, ~2 min): OSSL_Validation_Methods_WIN.zip
The one idea: a validation method is not a formula for accuracy โ€” it is a protocol for hiding data. Each of the five hides data differently, so each answers a different question. Same data, same model, different hiding rule โ†’ different number.

Setup โ€” what we're working with

400 soil samples from the Open Soil Spectral Library (OSSL) v1.2. Each sample is a reflectance spectrum, and we predict soil organic carbon (SOC) from it.

samples 400 bands 1051 (400โ€“2500 nm) target log10(SOC %) SOC raw range 0.05โ€“62.37 % model PLS, 10 components

Drawn from 4 real libraries: LUCAS.SSL 242, KSSL.SSL 133, ICRAF.ISRIC 23, LUCAS.WOODWELL.SSL 2.

Spectra and target distribution
How to read this: left โ€” 12 of the 400 real spectra; each line is one soil's light fingerprint. Right โ€” SOC on a log scale, because raw SOC is heavily right-skewed (a few soils are very carbon-rich) and that skew wrecks regression.

All five on one page

Schematic of the five methods
How to read this: blue = training data, orange = test, green = validation. Every method below is just a different way of painting this picture.
All methods compared
How to read this: the first seven bars are all in the same range โ€” they are all "random" hiding rules. The last red bar is a fundamentally different question. Keep that gap in mind; section 7 explains it.

1 ยท Holdout Validation

Rule: split once (here 70/30), train on the big part, test on the small part, report that one number.

The trap: the answer depends on which samples happened to land in the test set. Five different random splits, identical model, identical data โ†’ Rยฒ from 0.694 to 0.796.
Holdout vs repeated holdout
How to read this: left โ€” five single splits; the bar heights are the verdicts from the same model. Right โ€” 20 splits, and the mean sits in the middle with a real spread around it.
1
You get an unbiased estimate โ€” nothing about the test rows leaked into training.
2
But with high variance: the number is a lottery draw. Report one split and you can't tell if the model is good or you got lucky.
Analogy: it's like judging a student on one randomly chosen exam question. The question is fair, but the verdict swings wildly depending on which question they got.

2 ยท Repeated Holdout Validation

Rule: do the same thing many times with different splits, then report the mean and the spread. Cheapest possible fix for Method 1's variance.

On our data, 20 repeats give Rยฒ = 0.723 ยฑ 0.036 (range 0.656โ€“0.786). Note the mean barely moved versus a single split โ€” what you gained is the error bar.

Analogy: instead of one exam question, you give the student 20 randomly drawn questions and average. The average is a fairer verdict, and the spread tells you how much the verdict depends on the draw.

3 ยท Bootstrap Validation

Rule: draw n rows with replacement โ€” some rows appear several times, and about 37% never appear at all. Train on the drawn rows, then score the ones you missed. Those leftovers are the out-of-bag (OOB) samples.

Bootstrap resampling and optimism
How to read this: left โ€” one resample of 400 rows; 143 rows were never drawn (those are your free test set). Middle โ€” the optimism problem: scoring a model on its own training rows (apparent error 0.2942) looks better than reality (OOB 0.3208). Right โ€” the spread across all 200 resamples.

The .632 estimator blends the two: 0.368ร—apparent + 0.632ร—OOB = 0.3110. The exact split of the weights comes from the bootstrap theory of optimism โ€” you don't need to memorise it, just know it exists to correct the too-rosy apparent error.

The subtlety: because rows are duplicated in the in-bag set, the model partly memorises them. That is exactly why you must score on the OOB rows and never on the training rows.

4 ยท 3-Way Holdout

Rule: three parts, three jobs โ€” train fits the model, validation picks the hyperparameters, test is touched once at the very end to report.

Here we tune one thing: how many PLS components to use (grid 1โ€“20).

Three-way holdout
How to read this: left โ€” error curves for all three sets; the validation curve picks ncomp = 15. Right โ€” if you instead let the test set pick its own best model (ncomp = 7), the reported error drops to 0.3230 instead of the honest 0.3591.

That 0.0361 RMSE gap is the contamination. It is not a bug โ€” it is what happens when the test set stops being a neutral examiner and starts being part of the design.

Analogy: practice exams (validation) vs the final exam (test). If the teacher writes the final after seeing which questions you got right, your final score flatters you.

5 ยท Cross-Validation (k-fold, and Leave-One-Out)

Rule: split into k folds; train on kโˆ’1, test on the one left out; rotate so every row is tested exactly once; average the k scores.

k-fold CV
How to read this: left โ€” per-fold scores; the box height is the disagreement between folds. Right โ€” colour shows which fold tested each row, after sorting by SOC; every row is covered exactly once.
SchemeRยฒ (mean)sd across foldsRMSE (mean)
5-fold0.725 0.0520.3229
10-fold0.722 0.0870.3196
Leave-one-out (k = n = 400)0.737 โ€”0.3203
Counter-intuitive bit: more folds is not automatically better. LOO trains on nโˆ’1 rows every time, so its n folds are nearly identical โ€” the mean is good but the variance of the estimate can be understated. It also costs 400 model fits instead of 5. Note 10-fold here has a wider spread (0.087) than 5-fold (0.052) โ€” smaller test folds means noisier individual scores.
Leave-one-out diagnostics
How to read this: left โ€” predicted vs observed; right โ€” residuals. The far-right panel is the useful one: predictions go wrong at very low and very high SOC, which is a data-coverage problem, not a validation-method problem.

6 ยท All five side by side

MethodRยฒRMSENote
Holdout (single split)0.7440.3315one number, high variance
Repeated holdout ร—200.723โ€”sd Rยฒ 0.036
Bootstrap OOB0.7370.3208200 resamples
3-way (test, honest)0.7040.3591ncomp=15 from validation
5-fold CV0.7250.3229sd Rยฒ 0.052
10-fold CV0.7220.3196sd Rยฒ 0.087
Leave-one-out0.7370.3203400 fits
Leave-one-DATASET-out0.2030.4551the honest one

7 ยท The twist that matters for soil data

Every method above hides rows at random. But two soil samples from the same region are usually similar, so a random split can leave near-copies of your test samples inside the training set. Your score then partly measures memorisation of neighbours, not the ability to predict a new place.

The honest version holds out an entire source library: train on the others, predict a library the model has never seen.

Leakage gap
How to read this: left โ€” holding out a whole library gives 0.762 (KSSL), 0.064 (ICRAF), and -0.215 (LUCAS). The dashed line is what random 5-fold claimed. Right โ€” the gap: 0.725 vs 0.203.
Read this carefully: Rยฒ = 0 means "no better than always predicting the average". A negative Rยฒ means worse than that. LUCAS scored -0.215 โ€” the model trained on American and African soils actively mispredicts European soils. Random 5-fold hid that completely behind a score of 0.725.
The lesson: the leakage gap is 0.522 Rยฒ. The hiding rule must match the question you actually care about. "A new sample from the same places" and "a new place, new instrument, new lab" are very different claims โ€” and only one of them survives this test.

๐Ÿ“‹ Quick answer sheet

MethodHow it hides dataAnswers the questionWatch out for
Holdoutone random split"a new sample, same population" high variance โ€” one lucky draw
Repeated holdoutmany random splitssame, with an error bar cost scales with repeats
Bootstrapresample with replacement; score OOB same, reusing data efficientlyapparent error is way too optimistic
3-way holdoutseparate set for tuningsame, with honest tuning tuning on test contaminates the score
k-fold CVrotate k test foldssame, every row tested once fold spread; smaller folds are noisier
LOO (k = n)leave exactly one row outsame, maximum training data n fits; folds nearly identical โ†’ variance understated

๐Ÿง  Checklist before you report a score

1
Is it one split or a distribution? If one, report the spread too.
2
Was the test set ever used for decisions (tuning, feature selection, early stopping)? If yes, the score is contaminated.
3
Are you scoring on training rows anywhere? Remember the apparent-vs-OOB gap of 0.0266 RMSE.
4
Do your folds respect the structure of the data (region, site, instrument, year)? If not, you are measuring interpolation, not generalisation.
5
Does your validation rule match the claim you're making in the paper?