Interestana
Home/News/NVIDIA Releases Kumo Tabular Foundation Models
MarkTechPost••3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

NVIDIA Releases Kumo Tabular Foundation Models

NVIDIA has released Kumo Tabular, a new family of open tabular foundation models (TFMs) designed for classification and regression tasks. This innovative approach eliminates the need for traditional training, hyperparameter tuning, and feature engineering, allowing predictions on new data rows in a single forward pass. The Kumo Tabular family includes Small, Medium, and Large versions, with parameter counts ranging from approximately 28 million to 215 million. These models are accessible through NVIDIA's open-source structured-data-models (SDM) library, which is built for GPU-native operations and supports structured data foundation models and preprocessing. The weights for Kumo Tabular are distributed under the OpenMDW-1.1 license, which explicitly permits commercial use. The accompanying SDM code is licensed under Apache-2.0 and requires Python 3.11 or later, along with PyTorch 2.7 or later, and is optimized for CUDA GPUs.

The SDM library itself is a significant component, providing a unified interface for various structured data foundation models. Beyond Kumo Tabular, it integrates TabICLv2, Google's TabFM, and KumoRelational, which is specifically designed for multi-table data. All models within the library utilize a shared in-context learning interface built upon a TableTensor container. The library also streamlines preprocessing, ensembling, and the prediction of multiple classes. The architecture of Kumo Tabular is based on a Transformer, incorporating column, row, and in-context attention mechanisms, building upon concepts introduced in prior models like TabICL and TabPFN. The model's pipeline consists of three primary stages. The first stage, cell embedding, processes numerical and categorical values using learned Fourier features, with distinct weights for each data type and without requiring imputation for missing values. The second stage, row embedding, employs induced self-attention for column attention, ensuring that computational cost scales linearly with the number of rows. Row attention, utilizing rotary positions, is responsible for learning feature interactions. This stage also compresses each row using four learnable [CLS] tokens. The final stage, in-context learning, involves a Transformer that operates on the row embeddings. In this stage, context rows can attend to each other, while query rows are restricted to attending only to the context rows. Crucially, the context rows do not have access to the query rows, allowing their keys and values to be computed once and reused efficiently. The model's output head generates class probabilities for classification tasks or 999 quantiles for regression tasks, providing a comprehensive point prediction.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next