← All posts / Models

Gradient Boosting's Worst Week: Prior Labs' TabPFN-3.5 Sweeps Every Tabular Leaderboard

Prior Labs' TabPFN-3.5 takes first place on both TabArena and BeyondArena, roughly 150 Elo ahead on messy real-world data — and ships a six-times-faster Fast variant in alpha.

Gradient Boosting's Worst Week: Prior Labs' TabPFN-3.5 Sweeps Every Tabular Leaderboard

While the frontier labs argue about pacing and trillion-dollar funding rounds, the least glamorous format in machine learning — the spreadsheet — just had a frontier moment of its own. On September 15, 2026, Freiburg- and Oxford-spun Prior Labs released TabPFN-3.5, the newest version of its tabular foundation model, and it arrives with a clean sweep: first place on TabArena, first place on BeyondArena, and a family of four variants that stretches from a six-times-faster latency tier to a “Thinking” mode that spends test-time compute the way a reasoning LLM does.

If you don’t follow tabular machine learning, it is worth pausing on why this matters. Structured rows and columns — insurance claims, transactions, sensor readings, CRM records — are the data format that actually pays the bills in industry, and for a decade that territory has belonged to gradient-boosted decision trees: XGBoost, LightGBM, CatBoost. Every Kaggle competition, every credit-scoring model, every churn model has been a tuning exercise on top of boosted trees. TabPFN-3.5 is the strongest evidence yet that the foundation-model playbook — pretrain once, predict in-context, no per-dataset training — can take that territory.

What TabPFN actually is

TabPFN is a transformer that was pretrained on a vast space of synthetic datasets, in the “prior-fitted network” tradition. At inference time you hand it your training table and your test rows, and it predicts in-context — no gradient descent on your data, no hyperparameter search, no feature pipeline roulette. The model has already “learned to learn” on millions of curved, noisy, heterogeneous synthetic problems, and it treats your dataset as just another prompt.

The lineage is short but fast-moving. TabPFN-2.5 arrived in November 2025 and scaled to 20 times the data cells of its predecessor. TabPFN-3, released in May 2026, introduced test-time compute scaling to tabular prediction and pushed the sample ceiling to 5,000 rows (older versions capped at 1,000). TabPFN-3.5 is the third major revision in ten months, and it targets the specific place where the previous versions were weakest: data that violates the clean IID assumption.

What’s new in 3.5

The headline claim from the technical report is precision itself: TabPFN-3.5 ranks first on both TabArena and BeyondArena. TabArena is the living IID benchmark — 51 curated datasets, more than 27 competing methods, including over ten tabular foundation models. BeyondArena goes where standard benchmarks don’t: 142 datasets spanning high-dimensional, grouped, temporal, high-cardinality and text-rich data. On BeyondArena, TabPFN-3.5 finishes roughly 150 Elo points ahead of the previous overall leader, and on the non-large subset it leads on text-rich, high-cardinality and high-dimensional data by up to about 250 Elo over the strongest previous baseline. It matches the strongest baseline on grouped data. Coverage of the announcement also reports a 99 percent win rate over classic machine learning and 1,000-row predictions in 0.17 seconds.

The release is a family, not a single checkpoint:

  • TabPFN-3.5 (base) — the leaderboard-sweeping default.
  • TabPFN-3.5-Plus — adds enhanced processing to extract signal from text-rich datasets: product descriptions, customer reviews, internal notes. It ranks second on STRABLE, a benchmark of 108 tabular datasets full of messy strings — behind only its sibling below.
  • TabPFN-3.5-Thinking — the accuracy ceiling, trading compute for results: 44 additional Elo on TabArena and about 20 more on BeyondArena versus the base model. Thinking specifically improves on temporal and grouped data — the “trained on one store, predict sales at another, using historic records to predict next month” regime that breaks most off-the-shelf models.
  • TabPFN-3.5-Fast (alpha) — the first latency-oriented variant, up to six times faster than the base model, aimed at production serving where milliseconds matter.

The through-line is a bet on where real data lives. “Most machine learning assumes rows are independent, features are clean and there is enough data to train a gradient boosting tree,” the report notes. “In reality, business data rarely fits this assumption.” Insurance tables mix claim amounts with reviewer conclusions; transactions link to merchants and customers; industrial sensors collect hundreds of measurements that shift by site and season. Those are exactly the regimes where 3.5 claims — and on two independent arenas, demonstrates — its largest gains.

Enterprise traction arrived before the benchmark sweep

What makes this release more than an academic leaderboard story is the customer list Prior Labs has assembled. Marshmallow, the insurance unicorn, reports that TabPFN “outperforms gradient boosting on real-life insurance datasets without any tuning” — and highlights prediction confidence intervals, which risk teams care about as much as point estimates. Hitachi is applying TabPFN to predictive maintenance in rail. Creditplus Bank uses it for car-loan approval. TD has explored it for financial forecasting at enterprise scale, BostonGene used it to identify immune-system profiles, Exito for media-spend forecasting, and Oxford Cancer Analytics partnered with Prior Labs on liquid biopsy and clinical decision support in lung disease.

Distribution is following. The 3.5 family is available through Prior Labs’ own API and MCP server (which now defaults to 3.5-Plus), on AWS SageMaker (Plus and Thinking), and in SAP AI Core (Plus) — a placement that puts a foundation model inside the ERP estate where most of the world’s messy tables actually live. Azure ML availability is promised “shortly.” Until September 29, 2026, all API and MCP users get 50 percent off standard TabPFN-3.5 token rates. The open-source pip package ships the base model and the Fast checkpoint, and SAP and AWS listings confirm enterprise-grade packaging was ready on day one.

The honest caveats

Three qualifications keep the hype in check. First, licensing: the TabPFN weights are released under a non-commercial license (the TabPFN-3 license covers the 2.5, 2.6 and 3 generations), so “open” here means open code and research-capable checkpoints, not a CatBoost replacement you can ship in any commercial product without talking to Prior Labs. Second, scale: the in-context sweet spot remains thousands of rows, not millions — for the largest datasets Prior Labs is still working with customers on the TabPFN-2.5 Scaling Mode approach, which is a different workflow from simply pointing the model at a warehouse. Third, these are vendor-run leaderboard results. The arenas are well-curated and the Elo gaps are large, but independent replications at the same scale are just starting to circulate.

The deeper counterpoint is that boosted trees are cheap, boring, interpretable and infinitely debuggable, and a decade of MLOps tooling assumes them. Foundation models for tables win on zero-shot accuracy and speed-to-first-prediction; trees win on unit economics at scale and institutional comfort. The interesting phase begins when those trade-offs collide inside one platform — which is precisely what the SAP and AWS listings represent.

Why this week matters

Tabular prediction is the largest by-value segment of applied ML, and it has been the segment most insulated from the foundation-model revolution. TabPFN-3.5’s double first place, plus a six-times-faster serving tier and an accuracy-maximizing Thinking mode, is the strongest signal yet that the insulation is failing. For data teams, the pragmatic read is simple: benchmark it against your tuned gradient boosting baseline this quarter — the zero-shot, no-tuning setup means the evaluation costs a day, not a sprint. For the industry, it is one more domain where “train a model per problem” is giving way to “prompt a model that already knows how to learn.” The frontier labs are arguing about who slows down; in the spreadsheet, someone just sped up.

All evaluation details — protocols, datasets, configurations and runtime environment — are documented in the technical report linked below.