TabFM-Auto · Technical Report

TabFM-Auto: Self-Evolving Pipelines for Tabular Foundation Models

Pairing a language-model coding agent's domain reasoning with a frozen 400M-parameter TabFM backbone's in-context numerical prior reaches 2013.0 Elo (#1) across all 51 TabArena datasets and transfers zero-shot across frozen tabular foundation models.

Figure 1 · Benchmark Overview

Overall TabArena Leaderboard (51 Datasets)

TabFM-Auto (Evolved Pipeline + Frozen TabFM) TabFM / TabFM+ (Frozen Backbone) Other Tabular Foundation Models AutoML Ensembles (4h) Unconstrained Coding Agent (Antigravity + Gemini 3.8 Flash) Tuned Baselines
Figure 1: TabArena leaderboard across all 51 datasets. Hover or click a bar for details.

01Why Pair a Coding Agent with a Frozen Tabular Foundation Model?

Frozen TFMs and coding agents fail in opposite ways. TabFM-Auto keeps the TFM frozen and lets the agent rewrite only the data pipeline, so every CV gain comes from the pipeline itself.

Paradigm A · Frozen Backbone

Plain TabFM

  • +Strong numerical prior: single forward pass with frozen weights θ*.
  • +Deterministic: zero retraining noise.
  • −Schema-blind: ignores column names, domain units, and formulas.
  • −Single-table limit: cannot parse auxiliary waveform/3D files.
TabArena (51) 1785.3 Elo · Baseline
Paradigm B · End-to-End Agent

Unconstrained Coding Agent

  • +Reads metadata: infers domain ratios and multi-file parsers.
  • +Expressive Python: writes custom cleaning & feature logic.
  • −Noisy joint search: fitting GBDTs/NNs masks small feature gains.
  • −CV overfitting: stochastic model tuning overfits validation splits.
TabArena (51) 1468.8 Elo · −316.5
Best of Both Worlds · Ours

TabFM-Auto (Agent + Frozen TabFM)

  • ✓Semantic synthesis: writes domain formulas, clinical hierarchies & file parsers.
  • ✓Frozen 400M TabFM: fast forward-pass evaluation with zero gradient updates.
  • ✓Clean attribution: 100% of CV gain comes from the data pipeline itself.
  • ✓Zero-shot transfer: boosts TabICLv2 (+143 Elo), TabPFN-3 (+131 Elo), and EXAONE (+89 Elo).
TabArena (51) 2013.0 Elo · +227.7 (#1)

What Does the Agent Write? Three Kinds of Pipeline Edits

I
Domain features
Named columns → known formulas
17 / 51 datasets · engineer()
II
Statistical features
Anonymous columns → label-free statistics
34 / 51 datasets · engineer()
III
Context & calibration
Which rows TabFM sees; output rescaling
95% of runs · sample() postprocess()
Every dataset is I or II; III edits stack on top. Badges = test-error reduction vs. plain TabFM. Click a card to open its pipeline.py.

02Methodology: Closed-Loop Pipeline Evolution

Given training data Dtrain=(Xtrain,ytrain), unlabeled test inputs Xtest, and metadata M, we partition Dtrain into 3 internal cross-validation folds without touching the test split. The coding agent searches over executable Python pipelines P=(Φclean,Φfeat,Sctx,Ψpost)∈𝒫 around frozen weights θ* to maximize the mean 3-fold validation score Uval:

P*=arg maxP∈𝒫Uval(y^val,yval)
where, on each internal CV fold (tr = training part, val = held-out part):
(X~tr,y~tr,X~val)=Φfeat(Φclean(Xtr,ytr,Xval))1Clean & engineer preprocess() · engineer()C=Sctx(X~tr,y~tr)2Build context sample()y^val=Ψpost(fθ*(X~val∣C),ytr,X~val)3Predict & calibrate TabFM · postprocess()
(1)
Figure 2 · System Architecture & Interactive Stage Inspector

Closed-Loop Pipeline Search & 5-Stage Test-Time Dataflow

(a) Closed-Loop Agentic Search (3-Fold CV on Dtrain) Zero gradient updates · Strict test-set isolation
1. LLM Coding Agent Metadata M
Reads column names, units, and file schemas; proposes edits to pipeline.py and retains candidates that improve 3-fold CV score Uval.
2. Candidate Pipeline P pipeline.py
Implements four modular Python hooks (preprocess, engineer, sample, postprocess) plus declarative TABFM_KWARGS presets.
3. Sandboxed 3-Fold Judge Frozen θ* (400M)
Evaluates P across 3 internal folds of Dtrain in a single forward pass per view, then freezes the best pipeline P* for test inference.
(b) Inference Dataflow Through Selected Pipeline P* Around Frozen TabFM Click any of the 5 stages below to compare Before (P0) → After (P*)
1. preprocess() · Φclean
Clean & Target Map
Recodes sentinels (0 → NaN) and transforms skewed ytrain via g(y).
2. engineer() · Φfeat
Domain & Structural FE
Synthesizes domain ratios, clinical hierarchies, OOF rates, and multi-file features (95.1% of runs).
3. sample() · Sctx
Context Sampling
Builds stratified or minority-oversampled context views within 16,384 rows.
4. TABFM_KWARGS · fθ*
Frozen TabFM (400M)
Configures generic presets (norms, SVD, NNLS) and runs zero-shot ICL inference.
5. postprocess() · Ψpost
Calibrate & Inv-Map
Applies inverse map g−1 and log-odds prior alignment → 2013 Elo (#1).
Stage 2: engineer(X_train, y_train, X_test) — Semantic & Structural Feature Engineering (Φfeat)
Modified in 95.1% of TabArena runs (97 / 102)

Why it matters: Attention over raw numbers ignores column headers and external files; writing explicit domain formulas (Strouhal/Reynolds numbers, water-to-binder laws, ICD-9 hierarchies) or multi-file parsers spares TabFM from approximating nonlinear ratios in context.

Before · Identity Baseline P0 airfoil_self_noise · Test RMSE: 1.0712
# Raw wind-tunnel columns passed as-is:
# ["frequency", "angle-of-attack",
#  "chord-length", "free-stream-velocity",
#  "suction-side-displacement-thickness"]

def engineer(X_train, y_train, X_test):
    return X_train, X_test
After · Evolved Pipeline P* Test RMSE: 1.0712 → 0.9165 (−14.4%)
def _phys(df):
    f, c = df["frequency"], df["chord-length"]
    u, d = df["free-stream-velocity"], df["suction-side-displacement-thickness"]
    out = df.drop(columns=["frequency", "free-stream-velocity", "chord-length"]).copy()
    out["log_f"], out["log_d"], out["log_c"] = np.log10(f), np.log10(d), np.log10(c)
    out["log_St_d"], out["log_St_c"] = np.log10(f * d / u), np.log10(f * c / u)
    out["log_Re_c"], out["log_Re_d"] = np.log10(u * c / 1.5e-5), np.log10(u * d / 1.5e-5)
    out["log_d_over_c"] = np.log10(d / c)
    out["bpm_scale"] = 10.0 * np.log10(d) + 50.0 * np.log10(u / 340.46)
    return out

def engineer(X_train, y_train, X_test):
    return _phys(X_train), _phys(X_test)
Figure 2: Closed-loop search (a) and test-time dataflow (b).

03TabArena Pipeline Explorer (All 51 Datasets)

All 51 evolved pipelines. Pick a dataset and a stage to read its code, or open Write-up for notes.

Featured:
Open in full tab ↗

04TabArena Results, Taxonomy & Search Dynamics

Search runs once per dataset with 3-fold CV on fold 0's training split (up to 96 evaluations or 6 hours on one H100); the frozen P* is then scored on every official test fold. Over plain TabFM, TabFM-Auto gains +197.6 Elo on classification and +467.0 Elo on regression.

Table 1: Taxonomy of Evolved TabArena Pipelines Across All 51 Datasets

Category Representative Features Engineered in Code Datasets Mean Δ Error Median Δ Error
I. Domain-Knowledge FE Aeroacoustic Strouhal numbers (airfoil_self_noise), water-to-binder laws (concrete_strength), ICD-9 clinical hierarchies (Diabetes130US), SDSS color indices (SDSS17), DNA splice motifs (splice) 17 / 51 7.3% to 8.2% 3.0%
II. Statistical & Structural FE Transductive entity-graph degrees (Amazon_employee), missingness signatures (polish_bankruptcy), collinearity pruning & SVD projections (Bioresponse) 34 / 51 2.3% to 3.0% 1.7%
III. Context & Calibration Minority-oversampled context views in sample() and log-odds class-prior alignment in postprocess() (hiva_agnostic) 95% of runs 7.6% to 8.4% Log-Loss —
Table 1: Domain formulas on named schemas (I) cut test error nearly 3× more than statistical transforms on anonymized tables (II). Click a dataset to open its evolved pipeline.py.
Figure 3 · Search Dynamics & Generalization

Validation-to-Test Trajectory Across All 51 Datasets (0% → 100% Search Budget)

Checkpoint:
(a) 3-Fold Search Validation Gain (%)
(b) Official Test Error Reduction Across All Folds (%)
(c) Official Test TabArena Elo Progression
Figure 3: Search dynamics. CV gains (a) carry over to test error (b) and Elo (c); TabFM-Auto passes TabFM+ (1856 Elo) within the first 2.5% of the budget.

05Controlled Ablations & Zero-Shot Pipeline Transfer

Ablation 1: Frozen TabFM vs. Unconstrained Agent

Under the exact same agent (Antigravity + Gemini 3.8 Flash) and 6-hour budget, training GBDTs/PyTorch from scratch reaches 1468.8 Elo — 510.8 Elo below TabFM-Auto (1979.6).

Method Working Models Overall Cls. Reg.
Unconstrained Coding Agent Antigravity + Gemini 3.8 Flash GBDTs, PyTorch, Ensembles 1468.8 1472.0 1615.3
Plain TabFM TabFM (frozen) 1785.3 1768.7 2045.9
TabFM+ TabFM (frozen + presets) 1856.0 1836.7 2169.2
TabFM-Auto Antigravity + Gemini 3.8 Flash TabFM (frozen + evolved P*) 1979.6 1937.5 2347.8

Figure 4: Zero-Shot Transfer to Other Frozen TFMs

Pipelines evolved for TabFM, reused unchanged (no re-search) on three other frozen TFMs:

06Multi-File Autonomous Engineering on MLE-Bench-Tabular

On MLE-Bench-Tabular (8 Kaggle competitions, 12 external agents), much of the signal lives in auxiliary files: sensor waveforms, GNSS logs, and 3D molecular geometry. engineer() writes parsers for them (FFT bands, seismic STA/LTA triggers, Karplus angles, Kalman-smoothed GNSS fixes) and joins the results into one table for frozen TabFM.

Figure 5 · MLE-Bench-Tabular

MLE-Bench-Tabular Overall Elo Leaderboard (13 Agents)

Figure 5: MLE-Bench-Tabular. TabFM-Auto ranks #1 at 1827 Elo (+110 over CAIR MARS+). Click a competition card to open its pipeline below.

07MLE-Bench-Tabular Pipeline Explorer (All 8 Competitions)

All 8 evolved pipelines, with medal thresholds and write-ups.

Competitions:
Open in full tab ↗

BibTeXCitation

@article{fu2026tabfmauto,
  title   = {TabFM-Auto: Self-Evolving Pipelines for Tabular Foundation Models},
  author  = {Deqing Fu and Huangyuan Su and Rajat Sen and Taman Narayan and Sujay Sanghavi and Abhimanyu Das and Weihao Kong},
  year    = {2026},
  journal = {arXiv preprint arXiv:2609.37989},
  url     = {https://arxiv.org/abs/2609.37989}
}