A pre-registered comparison of 11 models and 6 preprocessing chains on the
Open Soil Spectral Library (OSSL v1.2) · computed on the UKY workstation
(Intel Core Ultra 9 285, 24 cores, NVIDIA RTX PRO 2000 Blackwell)
What this is. Three nutrient targets were predicted from soil spectra with every model
on identical data splits and identical preprocessing, so accuracy differences reflect the
algorithm, not the experimental setup. 111 configurations, 65 minutes of compute.
The headline: TabPFN won all three targets — and the gap is largest exactly where
prediction is hardest.
TabPFN
best on N, K and P
0.915
best R² (total N)
60,570
N samples used
111
configurations run
1. Framework
The design was fixed before any result was seen. Nothing was tuned on the test set.
Split. One stratified 70/30 split per target, seeded once and shared by
every model and preprocessing. Identical train rows, identical test rows.
Experiment A — models. 11 algorithms, all on snv_sg input.
Experiment B — preprocessing. 6 chains × 4 models, same split.
Experiment C — cross-library transfer. Train on one library, test on another.
This is the honest test of generalisation, and it is where everything falls apart.
Metrics. R², RMSE, MAE, bias, plus RPD and RPIQ because RPD assumes a normal
target and extractable P and exchangeable K are both strongly right-skewed.
Interpretability. SHAP and LIME — reported separately, labelled for what they
are. They explain a fitted model; they are not competitors in accuracy.
SNV removes multiplicative light-scatter effects. Derivatives remove additive baseline
drift and sharpen overlapping absorption bands. MSC rescales each spectrum against a
reference.
2. The data — and one finding that changed the design
Targets were joined from OSSL soillab to spectra on
id.layer_uuid_txt. Three things had to be established before modelling, and
each one is a real property of the release rather than a nuisance to work around:
Target
Method used
Spectra
n
Libraries
Total N
n.tot_usda.a623 (KSSL), n.tot_iso.11261 (LUCAS)
vis-NIR, 1051 bands, 400–2500 nm
60,570
LUCAS 40,175 · KSSL 19,806 · Woodwell 589
Exchangeable K
k.ext_usda.a725 (cmolc/kg)
vis-NIR, 1051 bands
44,489
LUCAS 40,175 · ICRAF 3,674 · KSSL 51
Extractable P
Mehlich-3 / ISO / Bray-1
MIR, 1701 bands, 600–4000 cm⁻¹
36,264
KSSL 11,155 · AFSIS1 · Garrett
Finding 1 — P is invisible to vis-NIR. Every OSSL layer carrying an extractable-P
value has no vis-NIR spectrum in this release; 27,011 P values exist and
zero of them overlap a complete vis-NIR scan. P had to be modelled from MIR.
That is not a limitation of the analysis — it is a fact about the data, and it matters
for anyone planning a vis-NIR P campaign.
Finding 2 — the band matrix is sparse per library. Only 17.4% of joined rows have a
complete spectrum across the native 350–2500 nm grid, because libraries start at different
wavelengths (LUCAS at 400 nm) and several libraries (AFSIS1, CAF, Garrett, Schiedung,
AFSIS2, Serbia) carry no vis-NIR at all — 70,726 MIR-only rows. Restricting to
400–2500 nm and requiring a complete spectrum lifted usable rows to 64,244. Zero-filling
blanks would have injected a library-identifying artefact into every model.
3. Results — Experiment A: models
All models on snv_sg, identical split. Sorted by R².
Total N (vis-NIR) — n_test = 3,600
Model
R²
RMSE
RPD
RPIQ
Bias
s
TabPFN
0.911
0.1616
3.34
1.55
+0.0005
43.7
SVR
0.859
0.2031
2.66
1.23
-0.0064
18.5
Cubist
0.858
0.2037
2.65
1.23
-0.0150
58.2
XGBoost
0.807
0.2372
2.28
1.05
-0.0072
13.9
MLP
0.806
0.2378
2.27
1.05
-0.0359
24.8
LightGBM
0.805
0.2383
2.27
1.05
-0.0048
5.8
CatBoost
0.793
0.2458
2.20
1.02
-0.0076
12.6
RF
0.787
0.2492
2.17
1.00
-0.0019
91.5
CNN1D
0.769
0.2597
2.08
0.96
+0.0605
8.2
Ridge
0.752
0.2693
2.01
0.93
-0.0019
0.2
PLS
0.719
0.2863
1.89
0.87
-0.0033
14.7
TabPFN reaches R² = 0.911 (RPD 3.34). SVR and Cubist follow at 0.859, then the boosted
trees cluster at 0.79–0.81. PLS — still the chemometrics default — lands last at 0.719.
Exchangeable K (vis-NIR) — n_test = 3,610
Model
R²
RMSE
RPD
RPIQ
Bias
s
TabPFN
0.509
0.6604
1.43
1.33
+0.0220
42.6
SVR
0.381
0.7416
1.27
1.18
-0.0866
42.9
Cubist
0.270
0.8056
1.17
1.09
-0.0911
56.5
MLP
0.265
0.8081
1.17
1.09
+0.0071
20.9
XGBoost
0.249
0.8172
1.15
1.07
+0.0352
14.1
Ridge
0.248
0.8175
1.15
1.07
+0.0384
0.2
CatBoost
0.244
0.8195
1.15
1.07
+0.0269
12.7
PLS
0.229
0.8276
1.14
1.06
+0.0405
0.9
CNN1D
0.212
0.8368
1.13
1.05
+0.0593
6.5
RF
0.206
0.84
1.12
1.05
+0.0710
74.9
LightGBM
0.199
0.8438
1.12
1.04
+0.0469
6.1
Much harder. TabPFN gets 0.509; everything else is at or below 0.381. No model clears
RPD 1.5, i.e. all are in the "poor" band for practical use.
Extractable P (MIR) — n_test = 3,600
Model
R²
RMSE
RPD
RPIQ
Bias
s
TabPFN
0.493
35.8
1.41
0.81
-1.2570
42.8
LightGBM
0.338
40.91
1.23
0.71
-0.4915
9.7
XGBoost
0.315
41.62
1.21
0.70
-0.0464
24.9
RF
0.291
42.34
1.19
0.69
+0.9854
126.2
MLP
0.290
42.37
1.19
0.69
-1.7631
27.6
CatBoost
0.284
42.54
1.18
0.68
-1.1537
19.9
Cubist
0.270
42.97
1.17
0.68
-5.0055
85.8
Ridge
0.237
43.92
1.15
0.66
-1.7503
0.4
SVR
0.187
45.34
1.11
0.64
-10.2419
66.3
PLS
0.173
45.72
1.10
0.64
-1.4645
1.5
CNN1D
0.106
47.54
1.06
0.61
-5.7356
6.3
TabPFN 0.493, then LightGBM 0.338 and XGBoost 0.315. Note CNN1D is worst on P (0.106)
despite MIR having the clearest physical signal.
Model ranking per target. TabPFN (red) leads all three. The dotted line at 0.5 is a reading aid, not a threshold.Full metrics for every model and target.Accuracy against wall-clock cost. TabPFN sits at the top-right of every target: highest accuracy at mid cost. LightGBM is the cheap option; RF is expensive for what it returns.
4. Results — Experiment B: preprocessing
Same split, four models, six chains. Cell shading marks the best chain within each
row.
Total N — R² by preprocessing and model
preproc
PLS
RF
XGBoost
TabPFN
raw
0.691
0.774
0.781
0.844
snv
0.721
0.787
0.809
0.910
snv_sg
0.719
0.787
0.807
0.911
snv_d1
0.754
0.851
0.877
0.915
msc
0.702
0.777
0.797
0.910
msc_sg
0.699
0.779
0.793
0.910
Exchangeable K
preproc
PLS
RF
XGBoost
TabPFN
raw
0.241
0.106
0.074
0.393
snv
0.229
0.216
0.250
0.511
snv_sg
0.229
0.206
0.249
0.509
snv_d1
0.231
0.065
0.397
0.612
msc
-0.474
0.175
0.126
0.513
msc_sg
-0.519
0.176
0.125
0.511
Extractable P (MIR)
preproc
PLS
RF
XGBoost
TabPFN
raw
0.170
0.215
0.231
0.461
snv
0.175
0.294
0.327
0.492
snv_sg
0.173
0.291
0.315
0.493
snv_d1
0.214
0.365
0.427
0.536
msc
0.174
0.288
0.333
0.495
msc_sg
0.173
0.286
0.325
0.496
Finding 3 — the first derivative (snv_d1) is the best single preprocessing choice on all
three targets, for every model. On K it is dramatic: XGBoost goes 0.249 → 0.397 and
TabPFN 0.509 → 0.612. Derivative preprocessing sharpens the overlapping absorption features
that scatter correction alone leaves smeared. On P, MSC is outright harmful to PLS
(R² = −0.47, worse than predicting the mean) while helping tree models.
Preprocessing effect. The derivative chain (snv_d1) is the rightmost point of every curve and the highest for every model.
5. Results — Experiment C: cross-library transfer
This is the test that matters for deployment and it is the one that fails. A model is
trained on one spectral library and applied to another. Negative R² means worse than
predicting the test set's own mean.
Target
Train library
Test library
R²
RMSE
RPD
n train
n test
N
KSSL.SSL
LUCAS.SSL
-0.184
0.4068
0.919
3,924
7,959
N
LUCAS.SSL
KSSL.SSL
0.443
0.556
1.340
7,959
3,924
K
ICRAF.ISRIC
LUCAS.SSL
-0.189
1.239
0.917
991
10,836
K
LUCAS.SSL
ICRAF.ISRIC
-29.145
2.934
0.182
10,836
991
P
AFSIS1.SSL
KSSL.SSL
-1.035
66.24
0.701
601
11,155
P
KSSL.SSL
AFSIS1.SSL
-0.474
31.78
0.824
11,155
601
Finding 4 — the models do not transfer across libraries. KSSL→LUCAS scores
R² = −0.184 on N. On K, LUCAS→ICRAF scores R² = −29.1. Every transfer is negative or
weak. Meanwhile random 5-fold within a library looks healthy. This is the single most
important caveat in the whole study: a good within-library R² tells you nothing about
whether the model works on someone else's soil, instrument, or lab. Anyone reporting only
random-split accuracy on OSSL is reporting an optimistic number.
Cross-library transfer. Bars below zero mean the model is worse than the mean. Only one transfers usefully at all (LUCAS→KSSL, N, R²=0.44).
6. Interpretability — SHAP and LIME
These are not models. SHAP and LIME explain a fitted model's behaviour. They have no
accuracy of their own and are not competitors in any of the tables above. They are included
because "where does the model look?" is a different and equally important question from
"how accurate is it?".
N concentrates on ~1408–1414 nm and 1882 nm, plus visible bands near 400 and
670 nm. The 1400 nm region is an OH/H₂O overtone and 1880 nm a water band; the visible
bands are the humus/iron colour region that carries organic-matter information. Soil total
N is largely bound in organic matter, so this is physically sensible.
K lands on 1654, 1696, and 2496–2500 nm. Reported K absorption features in
vis-NIR sit near 1400 and 1450 nm, and much of the 2450–2500 nm signal is clay
(Al-OH) — clay is a major exchange-site host, so the model is plausibly picking up the
exchange complex rather than K itself. That distinction is a hypothesis, not a claim.
P (MIR) is dominated by 3628–3698 cm⁻¹ and 1554–1558 cm⁻¹, with 600 cm⁻¹ also
salient. Note the model did not centre on the ~1030–1100 cm⁻¹ P-O region that
would be the textbook expectation — consistent with the low accuracy and with P being
largely present as occluded/sorbed species rather than free phosphate.
SHAP importance across the spectrum, XGBoost. Dashed lines mark textbook absorption positions. Y-axes are independent — magnitudes are not comparable between panels.
Do the predictions actually track the measurements?
SHAP says where a model looks; parity says whether the numbers are right. Below,
TabPFN's predicted values against measured values on the held-out test set. Points on the
dashed 1:1 line are perfect predictions.
Parity plots, TabPFN on the test set. N tracks the 1:1 line closely across the whole range. K and P compress toward the mean at the extremes — the classic signature of a model that has run out of signal, and the visual counterpart of their low RPD.
7. Pros, cons and when to use each model
Model
Advantages
Limitations
Partial Least Squares linear, latent variables
The workhorse of chemometrics: compresses 1051 collinear bands into a few latent components. Fast, interpretable, decades of precedent.
Captures only linear structure; needs the right component count
Ridge regression linear, L2
Extremely fast baseline. Useful as a floor for what linear structure alone buys.
Underfits the nonlinear relationship between spectra and nutrients
Random Forest bagged trees
Robust, hard to break, handles collinearity well, gives feature importance free.
Memory-hungry; extrapolates poorly; can be beaten by boosting
Support Vector Regression (RBF) kernel method
Excellent on high-dimensional, moderate-n spectral data. Strong here.
Needs a big kernel cache at this width; tuning C/epsilon matters; no feature importance
Multi-layer perceptron neural net, dense
Learns nonlinear interactions; scales to big n.
No spatial/spectral inductive bias, so it needs more data than a CNN; opaque
Gradient-boosted trees boosting
Strong general-purpose tabular default; handles collinearity and outliers; fast.
No extrapolation beyond training range; many hyperparameters
LightGBM boosting, histogram
Fastest of the boosted trees here (5.8s on N vs 92s for RF).
Leaf-wise growth can overfit small n; needs regularisation care
Slower to train than LightGBM; not better on these targets
Cubist rule-based M5 trees
The classic soil-spectroscopy algorithm: piecewise-linear rules suit local spectral behaviour. Very strong on N here.
Commercial-provenance model; rule count is a tuning knob; slower (58s)
TabPFN (tabular foundation model) in-context transformer
Pretrained transformer that does Bayesian inference in a single forward pass. Best on all three targets and notably better than everything else. No hyperparameter search needed.
Trained on synthetic priors, so its uncertainty is approximate; needs a license token; slower than boosting; capped here at 8000 training rows
best in P
1-D convolutional neural net deep learning
Learns its own spectral filters, so it can find local absorption features that fixed preprocessing might destroy. Trains on raw spectra.
Needs more data and tuning to shine; worst or near-worst on all three targets here
8. What actually follows from this
Use TabPFN first. It won every target without hyperparameter tuning, and its
margin grew as the task got harder (N: +0.05 R² over the runner-up; K: +0.13; P: +0.16).
This agrees with Barkov et al. (2026), who found TabPFN best across 85 pedometric
regression tasks, and with Huang et al. (2026) on MIR.
The marginal returns are in preprocessing, not architecture. Switching from raw
to snv_d1 bought XGBoost +0.09 R² on N and +0.15 on K — comparable to the gain from
switching algorithms entirely.
Report cross-library results. Within-library accuracy overstates field
performance badly. The transfer experiment is the honest one.
Manage expectations per element. N is predictable (RPD 3.4), P and K are not
(RPD 1.4–1.5). That ordering matches the literature: total N is bound to organic matter and
has real spectral expression; extractable P and K have weak or indirect responses.
If you need P, use MIR and expect a weak model. There is no vis-NIR route in
OSSL, and MIR only reaches R² ≈ 0.5. The literature reports the same difficulty.
Treat SHAP as hypothesis generation. It tells you where to look, not what the
chemistry is.
9. Caveats to state plainly
Cap of 12,000 rows per target. Chosen so the comparison finished in one session
and so the file could ship over the slow link. Absolute R² would likely rise with the full
sets; the ranking is the result, and rankings are robust to this at these n.
TabPFN was capped at 8,000 training rows (it is designed for "small data", and
this keeps memory bounded). Every other model saw all 8,400. So TabPFN's win is not an
artefact of easier data — it won on less.
P mixes lab methods (Mehlich-3, ISO, Bray-1) because no single method has enough
coverage with MIR. Method differences add noise to the target. The P number is therefore a
lower bound on what a single-method dataset could achieve.
CNN1D got one architecture and 40 epochs. A tuned CNN would likely do better;
these numbers are not a verdict on deep learning for spectra, only on this untuned
configuration.
Hyperparameters were reasonable defaults, not tuned per model. Tree models in
particular would gain from a search. TabPFN needs none, which is part of its appeal.
Single split, not repeated CV. Split uncertainty is not quantified here. The
earlier validation-methods work on the same library showed ±0.03 R² spread across repeated
holdouts, so differences below ~0.03 should not be over-read.
10. Reproducing this
Scripts and data are on the workstation at ~/npk/. Everything ran there;
Hermes drove it over the reverse SSH tunnel.
npk_compare.py # the whole experiment (fixed split, 3 experiments)
make_figs.py # the seven figures
npk_bench_small.npz # N and K: 12,000 rows x 1051 vis-NIR bands
npk_p_mir.npz # P: 12,000 rows x 1701 MIR bands
results_models.csv / results_preproc.csv / results_transfer.csv
11. References
Hollmann, N., Müller, S., Purucker, L., Krishnakumar, A., Körfer, M., Hoo, S.B.,
Schirrmeister, R.T., Hutter, F. (2025). Accurate predictions on small data with a tabular
foundation model. Nature 637, 319–326.
doi:10.1038/s41586-024-08328-6
Barkov, V., Schmidinger, J., Gebbers, R., Atzmueller, M. (2026). From field-scale to
large-scale spectral libraries: Tabular foundation models in soil spectroscopy.
arXiv:2608.00608. arxiv.org/abs/2608.00608
Huang, Y.-C., Minasny, B., et al. (2026). Zero-shot inference with Tabular Prior-data
Fitted Network (TabPFN) for soil MIR spectral analysis. Geoderma 471, 117880.
doi:10.1016/j.geoderma.2026.117880
Schmidinger, J., et al. (2026). A case for kriging-based spatial features with TabPFN in
digital soil mapping. Computers and Electronics in Agriculture.
S0168169925014589
Safanelli, J.L., Hengl, T., Parente, L.L., Minarik, R., Bloom, D.E., Todd-Brown, K.,
Gholizadeh, A., Mendes, W.S., Sanderman, J. (2025). Open Soil Spectral Library (OSSL):
Building reproducible soil calibration models through open development and community
engagement. PLoS ONE 20(1), e0296545.
doi:10.1371/journal.pone.0296545
Chang, C.-W., Laird, D.A., Mausbach, M.J., Hurburgh, C.R. (2001). Near-infrared
reflectance spectroscopy–principal components regression analyses of soil properties.
Soil Science Society of America Journal 65, 480–490. (RPD interpretation bands.)
Gozukara, G., et al. (2025). Prediction accuracy of pXRF, MIR, and Vis-NIR spectra for
soil properties — a review. Soil Science Society of America Journal 89, e0028.
Morellos, A., et al. (2016). Machine learning based prediction of soil total nitrogen,
organic carbon and moisture content by using VIS-NIR spectroscopy. Biosystems
Engineering 152, 104–116.
doi:10.1016/j.biosystemseng.2016.04.018
Mokere, R., et al. (2026). Soil spectroscopy improves mid-infrared soil property
prediction. Frontiers in Soil Science.
doi:10.3389/fsoil.2026.1760011
(reports unreliable spectroscopic prediction of soil phosphorus.)
Grinsztajn, L., Oyallon, E., Varoquaux, G. (2022). Why do tree-based models still
outperform deep learning on typical tabular data? NeurIPS Datasets and Benchmarks.
Lundberg, S.M., Lee, S.-I. (2017). A unified approach to interpreting model predictions.
NeurIPS. (SHAP.)
Ribeiro, M.T., Singh, S., Guestrin, C. (2016). "Why should I trust you?" Explaining the
predictions of any classifier. KDD. (LIME.)
Quinlan, J.R. (1992). Learning with continuous classes. 5th Australian Joint
Conference on Artificial Intelligence. (Cubist.)
Generated on the UKY workstation via the Hermes
reverse-SSH link · 111 configurations · 65 minutes · all numbers from
results_*.csv, not hand-entered.