TabPFN and TabICL against tuned XGBoost: the model that does not train won on fourteen tables out of fourteen
Pangram verdict · v3.3
We believe this text is mainly AI, with some human-written content.
AI likelihood · overall
AIArticle text · 1,375 words · 1 segments analyzed
Playing summary The claim has been going around for months and it is concrete enough to be measurable: a tabular foundation model predicts on a table without ever having trained on it and still beats tuned boosting. If that is true, half a decade of practice changes shape. Searching hyperparameters stops being a mandatory step and becomes a luxury that sometimes does not pay off. So I put it to the test on my own card, with fourteen datasets, four contenders and the same stopwatch for everyone. Your browser does not support HTML5 video.In 33 seconds with narration: two bots compete over a table. The one-eyed one looks and answers; the geared one tries twenty-five combinations before replying. Every number is a measured one. Muted by default: turn it on in the controls.Watch it in the reel viewer → What exactly these models do A tabular foundation model is pretrained on millions of synthetic tables generated on purpose. When a new table arrives, it does not adjust a single weight: it receives the training rows as context and produces the predictions in one forward pass. It is the same in-context learning idea we already know from language models, moved from words to columns. That is why the verb “train” sits oddly: the code still calls it fit, but inside there is no gradient descent, there is a copy of data to the card. That also explains why the cost shows up where you do not expect it. Fitting is nearly free and prediction is what pays, exactly the reverse of a tree. Put that way it sounds abstract, so here is the full journey of one row: a request arrives with an empty cell, the model weighs it against everything that already happened, resolves it in one pass and returns a future. Pick any of the six cases I measured to see it with their real data: Default riskSales campaignClinical readmissionRecidivismElectricity demandMolecular screeningThe requestThe contextA single passThe predictionA credit application arrives2 100 applications already settled · 10 columnsclf_num/credit.csvTabICL · nothing tuned · 0.8 sWill they repay?repaysdefaultsAUC 0.8528creditA contact enters the list2 100 calls already made · 7 columnsclf_num/bank-marketing.csvTabICL · nothing tuned · 0.8 sIs it worth calling them?signs updoes notAUC 0.8699bank-marketingA patient is discharged2 100 previous discharges · 7 columnsclf_num/Diabetes130US.csvTabICL · nothing tuned · 0.8 sWill they be readmitted?returnsdoes notAUC 0.6484Diabetes130USA case is assessed2 100 cases already closed · 11 columnsclf_cat/compas-two-years.csvTabICL · nothing tuned · 0.6 sWill they reoffend?reoffendsdoes notAUC 0.7329compas-two-yearsA market period closes2 100 previous periods · 7 columnsclf_num/electricity.csvTabICL · nothing tuned · 0.8 sDoes the price go up or down?updownAUC 0.8873electricityA candidate molecule arrives2 100 molecules already assayed · 419 columnsclf_num/Bioresponse.csvTabICL · nothing tuned · 6.0 sDoes it trigger a biological response?activeinertAUC 0.8667BioresponseThe requestThe contextA single passThe predictionA credit application arrives2 100 applications already settled · 10 columnsclf_num/credit.csvTabICL · nothing tuned · 0.8 sWill they repay?repaysdefaultsAUC 0.8528creditA contact enters the list2 100 calls already made · 7 columnsclf_num/bank-marketing.csvTabICL · nothing tuned · 0.8 sIs it worth calling them?signs updoes notAUC 0.8699bank-marketingA patient is discharged2 100 previous discharges · 7 columnsclf_num/Diabetes130US.csvTabICL · nothing tuned · 0.8 sWill they be readmitted?returnsdoes notAUC 0.6484Diabetes130USA case is assessed2 100 cases already closed · 11 columnsclf_cat/compas-two-years.csvTabICL · nothing tuned · 0.6 sWill they reoffend?reoffendsdoes notAUC 0.7329compas-two-yearsA market period closes2 100 previous periods · 7 columnsclf_num/electricity.csvTabICL · nothing tuned · 0.8 sDoes the price go up or down?updownAUC 0.8873electricityA candidate molecule arrives2 100 molecules already assayed · 419 columnsclf_num/Bioresponse.csvTabICL · nothing tuned · 6.0 sDoes it trigger a biological response?activeinertAUC 0.8667BioresponseDefault riskthe foundation model winsBanking’s most repeated case: deciding who gets lent to.datasetsizeTabICLTabPFNtuned XGBcredit3000 × 100.76670.75780.7533heloc3000 × 220.72220.73000.7078default-of-credit3000 × 200.69560.69670.6944All three datasets go to the foundation model. On heloc TabPFN takes 0.022 and TabICL 0.014: in credit risk that is not decoration.Context: clf_num/credit.csv from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.sha256 22a759296600c39a56884dafd84eb346be018c537ff096adc8c2ae7e0520f2a9Sales campaignthe foundation model winsWho to call first when there are a thousand contacts and time for a hundred.datasetsizeTabICLTabPFNtuned XGBbank-marketing3000 × 70.79440.79670.7833The only case where TabPFN ends up ahead of TabICL. Both beat tuned boosting.Context: clf_num/bank-marketing.csv from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.sha256 3433fecd416ce949692442882669881746e8a6b24b904a5f808cdcb196443e7eClinical readmissionthe foundation model winsA real hospital record: which patient gets readmitted.datasetsizeTabICLTabPFNtuned XGBDiabetes130US3000 × 70.59110.59330.5900On accuracy they nearly tie, but the area under the curve opens up sharply: 0.6484 against 0.6264. The risk ordering, which is what a triage uses, improves considerably more than accuracy suggests.Context: clf_num/Diabetes130US.csv from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.sha256 9be384ad7edbb9a98b509adb9cc08578e63bd87545b05d97d5147aa15b71385aRecidivismthe foundation model winsThe dataset that opened the debate on algorithmic bias in the courts.datasetsizeTabICLTabPFNtuned XGBcompas-two-years3000 × 110.67330.67670.6700The foundation model wins, and it is worth saying that a better model here does not make the use legitimate: this dataset’s argument was never about accuracy.Context: clf_cat/compas-two-years.csv from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.sha256 7e0bcef09ae633ad81e26b0a8f8f96dbc4ac3c29b9da0760d38ef2a75adde8d8Electricity demanda tieConsumption and price series from an electricity market.datasetsizeTabICLTabPFNtuned XGBelectricity3000 × 70.81780.81110.8178The tie. TabICL matches tuned boosting to the fourth decimal and TabPFN falls below. It is one of the two cases where the advantage does not show up.Context: clf_num/electricity.csv from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.sha256 d6005007c4b1f7ba88cf91cb1f231969cce3e49c98d445338c0ab0342ead5d7eMolecular screeningdepends on the model419 columns of chemical descriptors: the wide table.datasetsizeTabICLTabPFNtuned XGBBioresponse3000 × 4190.79220.75670.7744This is where TabPFN breaks: 0.7567 is worse than even untuned XGBoost, and it costs 18.1 seconds. TabICL holds up and comes first. The table’s width, not its length, is what squeezes.Context: clf_num/Bioresponse.csv from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.sha256 bd3d277821eb949df41219363f0386549019249776fccf7a9f08d3bec10b8727Pick a case above to follow one row’s journey. The context is the 2 100 training rows, the time is the one TabICL measured, and the orb’s arc draws that dataset’s area under the curve, from chance to a perfect score. What this is for, concretely The six cases in the diagram are not brochure examples: they are the datasets I measured with. But it is worth widening the map, because “tabular foundation model” sounds like a laboratory and the problem it solves is one of the most common there is. A table is a spreadsheet: rows that are cases and columns that are attributes. And the task is always the same, predict one column from the others: A customer table with tenure, plan, usage and complaints, to estimate which ones are going to leave next month. A transaction log with amount, merchant, hour and country, to flag which ones are fraud. A history of credit applications with income, debt and payment behaviour, to estimate who is not going to repay. Three of the datasets I measured are exactly that: credit, heloc and default-of-credit. Sensor readings from a machine, to anticipate when it is going to break. Patient records with symptoms and lab results, to prioritize who gets seen first. Diabetes130US, another of the datasets, is a real hospital record. That is half the data work done in any company. It is also the ground where deep learning had been losing for years: for tables, a good gradient-boosted tree ensemble was still the right answer, and the research that assembled these datasets was titled exactly that way, asking why trees still won. What changes if the promise holds Today, solving any of those cases has a ritual: prepare the data, choose a model, search hyperparameters, validate, repeat. The search is the boring part and the one that eats machine and human hours. In this benchmark, XGBoost’s search took up to 50 seconds per dataset; on a real problem with more rows and more combinations, it is minutes or hours. A tabular foundation model proposes skipping that ritual entirely. You hand it the table, it answers in a second, and that is that. No tree depth to choose, no learning rate, no cross-validation to decide among twenty-five candidates. Put in concrete terms: instead of spending the afternoon tuning a churn model, you have an answer in the time it takes to make coffee, and only then do you decide whether anything is worth refining. For exploring a table that just arrived, or for having an honest baseline before investing time, it is hard to beat.