01 What happened
NVIDIA has released Kumo-Tabular 1.0, a foundation model for classification and regression on tabular data. It uses labeled context rows and unlabeled query rows with the same columns to make predictions.
02 Key details
- According to NVIDIA, the model returns class probabilities or numeric predictions in a single forward pass, without task-specific training, gradient updates, or hyperparameter tuning.
- Pretraining used only synthetic tables generated from structural causal models; no real-world data was used.
- The release has three sizes: Small with about 28 million parameters, Medium with about 62 million, and Large with about 215 million. They were pretrained on about 35 million, 71 million, and 137 million synthetic tables, respectively.
- The model was evaluated on TabArena-v0.1 (51 datasets), TALENT (300), and BeyondArena, which uses random, temporal, and grouped splits. The suites were used only for evaluation; the score is undisclosed.
03 Why it matters
Kumo-Tabular gives data teams a way to make predictions from tables without training a separate model for each task. But NVIDIA has not published benchmark scores, so the model’s evaluation cannot be compared using reported numerical results.
04 Who it matters to
Analysts, machine learning engineers and teams working with tabular data.
Original sourceNVIDIA