TabFM-Auto: Self-Evolving Pipelines for Tabular Foundation Models
Pairing a language-model coding agent's domain reasoning with a frozen 400M-parameter TabFM backbone's in-context numerical prior reaches 2013.0 Elo (#1) across all 51 TabArena datasets and transfers zero-shot across frozen tabular foundation models.
Overall TabArena Leaderboard (51 Datasets)
| # | Method | Category | Bradley–Terry Elo ↑ | Dataset Wins ↑ | Improvability ↓ | G-Mean Error ↓ |
|---|
01Why Pair a Coding Agent with a Frozen Tabular Foundation Model?
Frozen TFMs and coding agents fail in opposite ways. TabFM-Auto keeps the TFM frozen and lets the agent rewrite only the data pipeline, so every CV gain comes from the pipeline itself.
Plain TabFM
- +Strong numerical prior: single forward pass with frozen weights θ*.
- +Deterministic: zero retraining noise.
- −Schema-blind: ignores column names, domain units, and formulas.
- −Single-table limit: cannot parse auxiliary waveform/3D files.
Unconstrained Coding Agent
- +Reads metadata: infers domain ratios and multi-file parsers.
- +Expressive Python: writes custom cleaning & feature logic.
- −Noisy joint search: fitting GBDTs/NNs masks small feature gains.
- −CV overfitting: stochastic model tuning overfits validation splits.
TabFM-Auto (Agent + Frozen TabFM)
- ✓Semantic synthesis: writes domain formulas, clinical hierarchies & file parsers.
- ✓Frozen 400M
TabFM: fast forward-pass evaluation with zero gradient updates. - ✓Clean attribution: 100% of CV gain comes from the data pipeline itself.
- ✓Zero-shot transfer: boosts
TabICLv2(+143 Elo),TabPFN-3(+131 Elo), andEXAONE(+89 Elo).
What Does the Agent Write? Three Kinds of Pipeline Edits
TabFM. Click a card to open its pipeline.py.
02Methodology: Closed-Loop Pipeline Evolution
Given training data , unlabeled test inputs , and metadata , we partition into 3 internal cross-validation folds without touching the test split. The coding agent searches over executable Python pipelines around frozen weights to maximize the mean 3-fold validation score :
preprocess() · engineer()2Build context sample()3Predict & calibrate TabFM · postprocess()Closed-Loop Pipeline Search & 5-Stage Test-Time Dataflow
pipeline.py and retains candidates that improve 3-fold CV score Uval.
pipeline.py
preprocess, engineer, sample, postprocess) plus declarative TABFM_KWARGS presets.
0 → NaN) and transforms skewed ytrain via g(y).Why it matters: Attention over raw numbers ignores column headers and external files; writing explicit domain formulas (Strouhal/Reynolds numbers, water-to-binder laws, ICD-9 hierarchies) or multi-file parsers spares TabFM from approximating nonlinear ratios in context.
# Raw wind-tunnel columns passed as-is:
# ["frequency", "angle-of-attack",
# "chord-length", "free-stream-velocity",
# "suction-side-displacement-thickness"]
def engineer(X_train, y_train, X_test):
return X_train, X_test
def _phys(df):
f, c = df["frequency"], df["chord-length"]
u, d = df["free-stream-velocity"], df["suction-side-displacement-thickness"]
out = df.drop(columns=["frequency", "free-stream-velocity", "chord-length"]).copy()
out["log_f"], out["log_d"], out["log_c"] = np.log10(f), np.log10(d), np.log10(c)
out["log_St_d"], out["log_St_c"] = np.log10(f * d / u), np.log10(f * c / u)
out["log_Re_c"], out["log_Re_d"] = np.log10(u * c / 1.5e-5), np.log10(u * d / 1.5e-5)
out["log_d_over_c"] = np.log10(d / c)
out["bpm_scale"] = 10.0 * np.log10(d) + 50.0 * np.log10(u / 340.46)
return out
def engineer(X_train, y_train, X_test):
return _phys(X_train), _phys(X_test)
03TabArena Pipeline Explorer (All 51 Datasets)
All 51 evolved pipelines. Pick a dataset and a stage to read its code, or open Write-up for notes.
04TabArena Results, Taxonomy & Search Dynamics
Search runs once per dataset with 3-fold CV on fold 0's training split (up to 96 evaluations or 6 hours on one H100); the frozen P* is then scored on every official test fold. Over plain TabFM, TabFM-Auto gains +197.6 Elo on classification and +467.0 Elo on regression.
Table 1: Taxonomy of Evolved TabArena Pipelines Across All 51 Datasets
| Category | Representative Features Engineered in Code | Datasets | Mean Δ Error | Median Δ Error |
|---|---|---|---|---|
| I. Domain-Knowledge FE | Aeroacoustic Strouhal numbers (airfoil_self_noise), water-to-binder laws (concrete_strength), ICD-9 clinical hierarchies (Diabetes130US), SDSS color indices (SDSS17), DNA splice motifs (splice) |
17 / 51 | 7.3% to 8.2% | 3.0% |
| II. Statistical & Structural FE | Transductive entity-graph degrees (Amazon_employee), missingness signatures (polish_bankruptcy), collinearity pruning & SVD projections (Bioresponse) |
34 / 51 | 2.3% to 3.0% | 1.7% |
| III. Context & Calibration | Minority-oversampled context views in sample() and log-odds class-prior alignment in postprocess() (hiva_agnostic) |
95% of runs | 7.6% to 8.4% Log-Loss | — |
pipeline.py.
Validation-to-Test Trajectory Across All 51 Datasets (0% → 100% Search Budget)
TabFM+ (1856 Elo) within the first 2.5% of the budget.
05Controlled Ablations & Zero-Shot Pipeline Transfer
Ablation 1: Frozen TabFM vs. Unconstrained Agent
Under the exact same agent (Antigravity + Gemini 3.8 Flash) and 6-hour budget, training GBDTs/PyTorch from scratch reaches 1468.8 Elo — 510.8 Elo below TabFM-Auto (1979.6).
| Method | Working Models | Overall | Cls. | Reg. |
|---|---|---|---|---|
| Unconstrained Coding Agent Antigravity + Gemini 3.8 Flash | GBDTs, PyTorch, Ensembles | 1468.8 | 1472.0 | 1615.3 |
| Plain TabFM | TabFM (frozen) | 1785.3 | 1768.7 | 2045.9 |
| TabFM+ | TabFM (frozen + presets) | 1856.0 | 1836.7 | 2169.2 |
| TabFM-Auto Antigravity + Gemini 3.8 Flash | TabFM (frozen + evolved P*) | 1979.6 | 1937.5 | 2347.8 |
Figure 4: Zero-Shot Transfer to Other Frozen TFMs
Pipelines evolved for TabFM, reused unchanged (no re-search) on three other frozen TFMs:
06Multi-File Autonomous Engineering on MLE-Bench-Tabular
On MLE-Bench-Tabular (8 Kaggle competitions, 12 external agents), much of the signal lives in auxiliary files: sensor waveforms, GNSS logs, and 3D molecular geometry. engineer() writes parsers for them (FFT bands, seismic STA/LTA triggers, Karplus angles, Kalman-smoothed GNSS fixes) and joins the results into one table for frozen TabFM.
MLE-Bench-Tabular Overall Elo Leaderboard (13 Agents)
07MLE-Bench-Tabular Pipeline Explorer (All 8 Competitions)
All 8 evolved pipelines, with medal thresholds and write-ups.
BibTeXCitation
@article{fu2026tabfmauto,
title = {TabFM-Auto: Self-Evolving Pipelines for Tabular Foundation Models},
author = {Deqing Fu and Huangyuan Su and Rajat Sen and Taman Narayan and Sujay Sanghavi and Abhimanyu Das and Weihao Kong},
year = {2026},
journal = {arXiv preprint arXiv:2609.37989},
url = {https://arxiv.org/abs/2609.37989}
}