Predicting soil NPK from spectra

A pre-registered comparison of 11 models and 6 preprocessing chains on the Open Soil Spectral Library (OSSL v1.2) · computed on the UKY workstation (Intel Core Ultra 9 285, 24 cores, NVIDIA RTX PRO 2000 Blackwell)

What this is. Three nutrient targets were predicted from soil spectra with every model on identical data splits and identical preprocessing, so accuracy differences reflect the algorithm, not the experimental setup. 111 configurations, 65 minutes of compute. The headline: TabPFN won all three targets — and the gap is largest exactly where prediction is hardest.
TabPFN
best on N, K and P
0.915
best R² (total N)
60,570
N samples used
111
configurations run

1. Framework

The design was fixed before any result was seen. Nothing was tuned on the test set.

  1. Split. One stratified 70/30 split per target, seeded once and shared by every model and preprocessing. Identical train rows, identical test rows.
  2. Experiment A — models. 11 algorithms, all on snv_sg input.
  3. Experiment B — preprocessing. 6 chains × 4 models, same split.
  4. Experiment C — cross-library transfer. Train on one library, test on another. This is the honest test of generalisation, and it is where everything falls apart.
  5. Metrics. R², RMSE, MAE, bias, plus RPD and RPIQ because RPD assumes a normal target and extractable P and exchangeable K are both strongly right-skewed.
  6. Interpretability. SHAP and LIME — reported separately, labelled for what they are. They explain a fitted model; they are not competitors in accuracy.

Preprocessing chains tested

raw snv snv_sg = SNV + Savitzky-Golay (w=11, p=2) snv_d1 = SNV + first derivative msc = multiplicative scatter correction msc_sg

SNV removes multiplicative light-scatter effects. Derivatives remove additive baseline drift and sharpen overlapping absorption bands. MSC rescales each spectrum against a reference.

2. The data — and one finding that changed the design

Targets were joined from OSSL soillab to spectra on id.layer_uuid_txt. Three things had to be established before modelling, and each one is a real property of the release rather than a nuisance to work around:

TargetMethod usedSpectranLibraries
Total Nn.tot_usda.a623 (KSSL), n.tot_iso.11261 (LUCAS) vis-NIR, 1051 bands, 400–2500 nm60,570 LUCAS 40,175 · KSSL 19,806 · Woodwell 589
Exchangeable Kk.ext_usda.a725 (cmolc/kg) vis-NIR, 1051 bands44,489 LUCAS 40,175 · ICRAF 3,674 · KSSL 51
Extractable PMehlich-3 / ISO / Bray-1 MIR, 1701 bands, 600–4000 cm⁻¹36,264 KSSL 11,155 · AFSIS1 · Garrett
Finding 1 — P is invisible to vis-NIR. Every OSSL layer carrying an extractable-P value has no vis-NIR spectrum in this release; 27,011 P values exist and zero of them overlap a complete vis-NIR scan. P had to be modelled from MIR. That is not a limitation of the analysis — it is a fact about the data, and it matters for anyone planning a vis-NIR P campaign.
Finding 2 — the band matrix is sparse per library. Only 17.4% of joined rows have a complete spectrum across the native 350–2500 nm grid, because libraries start at different wavelengths (LUCAS at 400 nm) and several libraries (AFSIS1, CAF, Garrett, Schiedung, AFSIS2, Serbia) carry no vis-NIR at all — 70,726 MIR-only rows. Restricting to 400–2500 nm and requiring a complete spectrum lifted usable rows to 64,244. Zero-filling blanks would have injected a library-identifying artefact into every model.

3. Results — Experiment A: models

All models on snv_sg, identical split. Sorted by R².

Total N (vis-NIR) — n_test = 3,600

ModelR²RMSERPDRPIQBiass
TabPFN0.9110.16163.341.55+0.000543.7
SVR0.8590.20312.661.23-0.006418.5
Cubist0.8580.20372.651.23-0.015058.2
XGBoost0.8070.23722.281.05-0.007213.9
MLP0.8060.23782.271.05-0.035924.8
LightGBM0.8050.23832.271.05-0.00485.8
CatBoost0.7930.24582.201.02-0.007612.6
RF0.7870.24922.171.00-0.001991.5
CNN1D0.7690.25972.080.96+0.06058.2
Ridge0.7520.26932.010.93-0.00190.2
PLS0.7190.28631.890.87-0.003314.7

TabPFN reaches R² = 0.911 (RPD 3.34). SVR and Cubist follow at 0.859, then the boosted trees cluster at 0.79–0.81. PLS — still the chemometrics default — lands last at 0.719.

Exchangeable K (vis-NIR) — n_test = 3,610

ModelR²RMSERPDRPIQBiass
TabPFN0.5090.66041.431.33+0.022042.6
SVR0.3810.74161.271.18-0.086642.9
Cubist0.2700.80561.171.09-0.091156.5
MLP0.2650.80811.171.09+0.007120.9
XGBoost0.2490.81721.151.07+0.035214.1
Ridge0.2480.81751.151.07+0.03840.2
CatBoost0.2440.81951.151.07+0.026912.7
PLS0.2290.82761.141.06+0.04050.9
CNN1D0.2120.83681.131.05+0.05936.5
RF0.2060.841.121.05+0.071074.9
LightGBM0.1990.84381.121.04+0.04696.1

Much harder. TabPFN gets 0.509; everything else is at or below 0.381. No model clears RPD 1.5, i.e. all are in the "poor" band for practical use.

Extractable P (MIR) — n_test = 3,600

ModelR²RMSERPDRPIQBiass
TabPFN0.49335.81.410.81-1.257042.8
LightGBM0.33840.911.230.71-0.49159.7
XGBoost0.31541.621.210.70-0.046424.9
RF0.29142.341.190.69+0.9854126.2
MLP0.29042.371.190.69-1.763127.6
CatBoost0.28442.541.180.68-1.153719.9
Cubist0.27042.971.170.68-5.005585.8
Ridge0.23743.921.150.66-1.75030.4
SVR0.18745.341.110.64-10.241966.3
PLS0.17345.721.100.64-1.46451.5
CNN1D0.10647.541.060.61-5.73566.3

TabPFN 0.493, then LightGBM 0.338 and XGBoost 0.315. Note CNN1D is worst on P (0.106) despite MIR having the clearest physical signal.

Model ranking per target. TabPFN (red) leads all three. The dotted line at 0.5 is a reading aid, not a threshold.
Model ranking per target. TabPFN (red) leads all three. The dotted line at 0.5 is a reading aid, not a threshold.
Full metrics for every model and target.
Full metrics for every model and target.
Accuracy against wall-clock cost. TabPFN sits at the top-right of every target: highest accuracy at mid cost. LightGBM is the cheap option; RF is expensive for what it returns.
Accuracy against wall-clock cost. TabPFN sits at the top-right of every target: highest accuracy at mid cost. LightGBM is the cheap option; RF is expensive for what it returns.

4. Results — Experiment B: preprocessing

Same split, four models, six chains. Cell shading marks the best chain within each row.

Total N — R² by preprocessing and model

preprocPLSRFXGBoostTabPFN
raw0.6910.7740.7810.844
snv0.7210.7870.8090.910
snv_sg0.7190.7870.8070.911
snv_d10.7540.8510.8770.915
msc0.7020.7770.7970.910
msc_sg0.6990.7790.7930.910

Exchangeable K

preprocPLSRFXGBoostTabPFN
raw0.2410.1060.0740.393
snv0.2290.2160.2500.511
snv_sg0.2290.2060.2490.509
snv_d10.2310.0650.3970.612
msc-0.4740.1750.1260.513
msc_sg-0.5190.1760.1250.511

Extractable P (MIR)

preprocPLSRFXGBoostTabPFN
raw0.1700.2150.2310.461
snv0.1750.2940.3270.492
snv_sg0.1730.2910.3150.493
snv_d10.2140.3650.4270.536
msc0.1740.2880.3330.495
msc_sg0.1730.2860.3250.496
Finding 3 — the first derivative (snv_d1) is the best single preprocessing choice on all three targets, for every model. On K it is dramatic: XGBoost goes 0.249 → 0.397 and TabPFN 0.509 → 0.612. Derivative preprocessing sharpens the overlapping absorption features that scatter correction alone leaves smeared. On P, MSC is outright harmful to PLS (R² = −0.47, worse than predicting the mean) while helping tree models.
Preprocessing effect. The derivative chain (snv_d1) is the rightmost point of every curve and the highest for every model.
Preprocessing effect. The derivative chain (snv_d1) is the rightmost point of every curve and the highest for every model.

5. Results — Experiment C: cross-library transfer

This is the test that matters for deployment and it is the one that fails. A model is trained on one spectral library and applied to another. Negative R² means worse than predicting the test set's own mean.

TargetTrain libraryTest libraryR²RMSE RPDn trainn test
NKSSL.SSLLUCAS.SSL-0.1840.40680.9193,9247,959
NLUCAS.SSLKSSL.SSL0.4430.5561.3407,9593,924
KICRAF.ISRICLUCAS.SSL-0.1891.2390.91799110,836
KLUCAS.SSLICRAF.ISRIC-29.1452.9340.18210,836991
PAFSIS1.SSLKSSL.SSL-1.03566.240.70160111,155
PKSSL.SSLAFSIS1.SSL-0.47431.780.82411,155601
Finding 4 — the models do not transfer across libraries. KSSL→LUCAS scores R² = −0.184 on N. On K, LUCAS→ICRAF scores R² = −29.1. Every transfer is negative or weak. Meanwhile random 5-fold within a library looks healthy. This is the single most important caveat in the whole study: a good within-library R² tells you nothing about whether the model works on someone else's soil, instrument, or lab. Anyone reporting only random-split accuracy on OSSL is reporting an optimistic number.
Cross-library transfer. Bars below zero mean the model is worse than the mean. Only one transfers usefully at all (LUCAS→KSSL, N, R²=0.44).
Cross-library transfer. Bars below zero mean the model is worse than the mean. Only one transfers usefully at all (LUCAS→KSSL, N, R²=0.44).

6. Interpretability — SHAP and LIME

These are not models. SHAP and LIME explain a fitted model's behaviour. They have no accuracy of their own and are not competitors in any of the tables above. They are included because "where does the model look?" is a different and equally important question from "how accurate is it?".
TargetTimeSamples explainedTop-10 bands
N0.1s4001410, 1408, 400, 1414, 1412, 1882, 670, 1406, 1404, 1386
K0.1s4001696, 1654, 400, 2500, 2498, 1650, 1700, 2496, 1710, 402
P0.1s4003660, 3662, 3652, 1556, 1558, 1554, 3674, 3646, 3654, 3698

Reading the profiles against known assignments:

SHAP importance across the spectrum, XGBoost. Dashed lines mark textbook absorption positions. Y-axes are independent — magnitudes are not comparable between panels.
SHAP importance across the spectrum, XGBoost. Dashed lines mark textbook absorption positions. Y-axes are independent — magnitudes are not comparable between panels.

Do the predictions actually track the measurements?

SHAP says where a model looks; parity says whether the numbers are right. Below, TabPFN's predicted values against measured values on the held-out test set. Points on the dashed 1:1 line are perfect predictions.

Parity plots, TabPFN on the test set. N tracks the 1:1 line closely across the whole range. K and P compress toward the mean at the extremes — the classic signature of a model that has run out of signal, and the visual counterpart of their low RPD.
Parity plots, TabPFN on the test set. N tracks the 1:1 line closely across the whole range. K and P compress toward the mean at the extremes — the classic signature of a model that has run out of signal, and the visual counterpart of their low RPD.

7. Pros, cons and when to use each model

ModelAdvantagesLimitations
Partial Least Squares
linear, latent variables
The workhorse of chemometrics: compresses 1051 collinear bands into a few latent components. Fast, interpretable, decades of precedent.Captures only linear structure; needs the right component count
Ridge regression
linear, L2
Extremely fast baseline. Useful as a floor for what linear structure alone buys.Underfits the nonlinear relationship between spectra and nutrients
Random Forest
bagged trees
Robust, hard to break, handles collinearity well, gives feature importance free.Memory-hungry; extrapolates poorly; can be beaten by boosting
Support Vector Regression (RBF)
kernel method
Excellent on high-dimensional, moderate-n spectral data. Strong here.Needs a big kernel cache at this width; tuning C/epsilon matters; no feature importance
Multi-layer perceptron
neural net, dense
Learns nonlinear interactions; scales to big n.No spatial/spectral inductive bias, so it needs more data than a CNN; opaque
Gradient-boosted trees
boosting
Strong general-purpose tabular default; handles collinearity and outliers; fast.No extrapolation beyond training range; many hyperparameters
LightGBM
boosting, histogram
Fastest of the boosted trees here (5.8s on N vs 92s for RF).Leaf-wise growth can overfit small n; needs regularisation care
CatBoost
boosting, ordered
Strong out-of-the-box, well-regularised, handles skewed targets.Slower to train than LightGBM; not better on these targets
Cubist
rule-based M5 trees
The classic soil-spectroscopy algorithm: piecewise-linear rules suit local spectral behaviour. Very strong on N here.Commercial-provenance model; rule count is a tuning knob; slower (58s)
TabPFN (tabular foundation model)
in-context transformer
Pretrained transformer that does Bayesian inference in a single forward pass. Best on all three targets and notably better than everything else. No hyperparameter search needed.Trained on synthetic priors, so its uncertainty is approximate; needs a license token; slower than boosting; capped here at 8000 training rowsbest in P
1-D convolutional neural net
deep learning
Learns its own spectral filters, so it can find local absorption features that fixed preprocessing might destroy. Trains on raw spectra.Needs more data and tuning to shine; worst or near-worst on all three targets here

8. What actually follows from this

  1. Use TabPFN first. It won every target without hyperparameter tuning, and its margin grew as the task got harder (N: +0.05 R² over the runner-up; K: +0.13; P: +0.16). This agrees with Barkov et al. (2026), who found TabPFN best across 85 pedometric regression tasks, and with Huang et al. (2026) on MIR.
  2. The marginal returns are in preprocessing, not architecture. Switching from raw to snv_d1 bought XGBoost +0.09 R² on N and +0.15 on K — comparable to the gain from switching algorithms entirely.
  3. Report cross-library results. Within-library accuracy overstates field performance badly. The transfer experiment is the honest one.
  4. Manage expectations per element. N is predictable (RPD 3.4), P and K are not (RPD 1.4–1.5). That ordering matches the literature: total N is bound to organic matter and has real spectral expression; extractable P and K have weak or indirect responses.
  5. If you need P, use MIR and expect a weak model. There is no vis-NIR route in OSSL, and MIR only reaches R² ≈ 0.5. The literature reports the same difficulty.
  6. Treat SHAP as hypothesis generation. It tells you where to look, not what the chemistry is.

9. Caveats to state plainly

10. Reproducing this

Scripts and data are on the workstation at ~/npk/. Everything ran there; Hermes drove it over the reverse SSH tunnel.

npk_compare.py       # the whole experiment (fixed split, 3 experiments)
make_figs.py         # the seven figures
npk_bench_small.npz  # N and K: 12,000 rows x 1051 vis-NIR bands
npk_p_mir.npz        # P: 12,000 rows x 1701 MIR bands
results_models.csv / results_preproc.csv / results_transfer.csv

11. References

  1. Hollmann, N., Müller, S., Purucker, L., Krishnakumar, A., Körfer, M., Hoo, S.B., Schirrmeister, R.T., Hutter, F. (2025). Accurate predictions on small data with a tabular foundation model. Nature 637, 319–326. doi:10.1038/s41586-024-08328-6
  2. Barkov, V., Schmidinger, J., Gebbers, R., Atzmueller, M. (2026). From field-scale to large-scale spectral libraries: Tabular foundation models in soil spectroscopy. arXiv:2608.00608. arxiv.org/abs/2608.00608
  3. Huang, Y.-C., Minasny, B., et al. (2026). Zero-shot inference with Tabular Prior-data Fitted Network (TabPFN) for soil MIR spectral analysis. Geoderma 471, 117880. doi:10.1016/j.geoderma.2026.117880
  4. Schmidinger, J., et al. (2026). A case for kriging-based spatial features with TabPFN in digital soil mapping. Computers and Electronics in Agriculture. S0168169925014589
  5. Safanelli, J.L., Hengl, T., Parente, L.L., Minarik, R., Bloom, D.E., Todd-Brown, K., Gholizadeh, A., Mendes, W.S., Sanderman, J. (2025). Open Soil Spectral Library (OSSL): Building reproducible soil calibration models through open development and community engagement. PLoS ONE 20(1), e0296545. doi:10.1371/journal.pone.0296545
  6. Chang, C.-W., Laird, D.A., Mausbach, M.J., Hurburgh, C.R. (2001). Near-infrared reflectance spectroscopy–principal components regression analyses of soil properties. Soil Science Society of America Journal 65, 480–490. (RPD interpretation bands.)
  7. Gozukara, G., et al. (2025). Prediction accuracy of pXRF, MIR, and Vis-NIR spectra for soil properties — a review. Soil Science Society of America Journal 89, e0028.
  8. Morellos, A., et al. (2016). Machine learning based prediction of soil total nitrogen, organic carbon and moisture content by using VIS-NIR spectroscopy. Biosystems Engineering 152, 104–116. doi:10.1016/j.biosystemseng.2016.04.018
  9. Mokere, R., et al. (2026). Soil spectroscopy improves mid-infrared soil property prediction. Frontiers in Soil Science. doi:10.3389/fsoil.2026.1760011 (reports unreliable spectroscopic prediction of soil phosphorus.)
  10. Grinsztajn, L., Oyallon, E., Varoquaux, G. (2022). Why do tree-based models still outperform deep learning on typical tabular data? NeurIPS Datasets and Benchmarks.
  11. Lundberg, S.M., Lee, S.-I. (2017). A unified approach to interpreting model predictions. NeurIPS. (SHAP.)
  12. Ribeiro, M.T., Singh, S., Guestrin, C. (2016). "Why should I trust you?" Explaining the predictions of any classifier. KDD. (LIME.)
  13. Quinlan, J.R. (1992). Learning with continuous classes. 5th Australian Joint Conference on Artificial Intelligence. (Cubist.)

Generated on the UKY workstation via the Hermes reverse-SSH link · 111 configurations · 65 minutes · all numbers from results_*.csv, not hand-entered.