MATH · IN · MODELS

Real TabPFNv2, TabICLv2, and Mitra build class-conditional prototype/vote geometry, causally confirmed via mechanism transplantation

measured in 1 paper

Biloš, Wilson, Schneider & Nevmyvaka (2026) study three real pretrained tabular foundation models: TabICLv2 reads out predictions via a nearest-prototype rule (per-class centroid of context-row activations), while TabPFNv2/Mitra use an attention-weighted vote at a specific layer; a linear probe under-explains the readout (0.859 acc) with only marginal gain from a bounding nonlinear-MLP probe [bilos-etal-2026-a-mechanistic-study-of-tabular-foundation-models] Causal interventions are decisive: forcing uniform attention drops TabPFNv2 accuracy from 0.87 to 0.49, transplanting one model's readout rule onto another's activations drops accuracy 30-40 percentage points, and zeroing TabPFNv2's positional-parameter matrix yields exact permutation invariance at no accuracy cost [bilos-etal-2026-a-mechanistic-study-of-tabular-foundation-models]

Method

Papers

A Mechanistic Study of Tabular Foundation Models — Biloš, Marin, Wilson, James T., Schneider, Anderson, Nevmyvaka, Yuriy2026 · arXiv:2605.21288