Intro

I have been working with models for tabular data a bit professionally and as a hobby. I did notice attempts over time to use neural nets to somehow beat decision tree based models. However, they all pretty much at most were on par with decision tree based models, even in regimes with millions of samples.

A recent post by Gaël Varoquaux caught my attention stating “Tabular foundation models have gone from a promising idea to a production-ready technology”. Notable, as he is one of the creators of scikit-learn - a tool I consider an essential work horse of data science.

Doing a quick search, I landed on a comparison from February 2026 that Gaël Varoquaux co-authored that even claims that tabular foundation models now “dethroned gradient-boosted trees at the top of predictive benchmarks”.

So interest was sufficiently piqued. Time to try and find out what has changed.

Roots

The origin of those referenced tabular foundation models seems to be the “Transformers Can Do Bayesian Inference” paper. In it, the type of model is referred to as Tabular Prior-Data Fitted Networks and abbreviated as TabPFN (See this part I, II and III of the mindfulmodeler’s blog post series for a more step-by-step explanation on how this type of model works.)

Luckily the same group that published the “Transformers Can Do Bayesian Inference” paper also published the very readable nanoTabPFN repo, alongside this paper, that provides the tools for the modelling, training and evaluation in a very concise form for trying this at home.

Trying this at home

The nanoTabPFN repo was sufficiently compact that I could migrate it to uv, for fun and reproducibility, and try and train models locally myself.

I wanted to see if this can really run locally on my machine as well as what sort of performance can be obtained in reasonable time. This resulted in nanoTabPFN-fork.

I’ve stuck with the experiment as designed in the original nanoTabPFN repo. This means training a binary classifier based on the NanoTabPFNModel class which sees only synthetic data during the training and evaluating the model against the same three openml datasets as in the paper, i.e. amazon employee access, blood transfusion service center and diabetes, without using that to guide the training.

The synthetic data originated in the first TabICL paper (TabICL is another flavor of TabPFN). It contains 300k datasets, 150 data points, 5 features, binary classification target as outlined on the nanoTabPFN wiki and apparently generated using the code referenced here.

Boldly kicking off the training with the defaults, not knowing what to expect, I found, to my surprise, that it is actually feasible to train nanoTabPFN models locally within reasonable time, no machine catching fire or the like.

Below the results of the training runs

Grid illustration of ROC AUC scores

Grid illustration of ROC AUC scores (y-axis) for a range of models (series) on three different data sets (Amazon employee access, blood transfusion service center, diabetes) over the number of training iterations of the nanoTabPFN model. The lines and shaded areas are mean and mean +- one standard deviation of training re-runs respectively.

display a blue wiggly line, the locally trained nanoTabPFN model over training iterations, with the ROC AUC score for three datasets and the results of reference models, the straight lines. The plot was generated using the experiment.ipynb notebook.

So to re-emphasize, the nanoTabPFN model performance, with all its wiggliness in the above plot is only measured on those datsets, it is not trained on them. It is trained on synthetic datasets. Still, its performance is partially equal or superior to other Machine Learning models that are explicitly trained on those three datasets.

Why the performance is so wiggly is a good question. The nanoTabPFN paper displays the original version of that plot in Figure 5 in Appendix B. There it seems more like an upwards trend. The reason for this discrepancy is unclear.

Also interesting to note. The original comparison in the nanoTabPFN paper was against TabPFN v2, since that was the most recent. Now including v3.5 it seems it is not necessarily better across the board.

Takeaway

A clear limitation of this type of model is the scaling (quadratic in both the number of samples and features used), limiting it currently to regimes of < 100k samples and a dozen or so features.

Similar to language models being used in a new use case, those tabular foundation models need an eval on the new use case they are supposed to be used in to verify performance, newer model versions don’t guarantee performance improvement.

However, the nanoTabPFN model, trained on synthetic data only, performed on par or better on the three previously unseen datasets than, e.g. a decision tree-based, models that were explicitly trained on those datasets!

Extrapolating this model behavior beyond the three datasets makes this approach definitely seem like one of those to keep an eye on, especially for low data regimes.

Going further

Both TabPFN and TabICL are also available as open source packages abiding by the scikit-learn API standard, making integration / migration super easy.

A range of TabICL tutorials from model interpretability to fine-tuning can be found here.

Happy adventuring!