Team Ai
Datasetpublic

Swapnil007-Curious/kanvas-gpu-failure

KANVAS GPU Task Failure Features Derived data for the KANVAS project: an interpretable Kolmogorov-Arnold Additive Model (KAAM) that predicts GPU cluster task failure. Code, notebooks and tests: https://github.com/Swapnil007-Curious/KANVAS Files File Rows Contents kanvas_feature_matrix.csv 18,890 10 scheduling and telemetry features plus the failure label, one row per task kanvas_phi_curves.csv 2,000 Learned per-feature curves phi_i(x) of the canonical… See the full description on the dataset page: https://huggingface.co/datasets/Swapnil007-Curious/kanvas-gpu-failure.

sourceHugging Facecc-by-4.0updated 11d agoView on Hugging Face
1likes53downloads
Dataset Card

KANVAS GPU Task Failure Features

Derived data for the KANVAS project: an interpretable Kolmogorov-Arnold Additive Model (KAAM) that predicts GPU cluster task failure. Code, notebooks and tests: https://github.com/Swapnil007-Curious/KANVAS

Files

FileRowsContents
kanvas_feature_matrix.csv18,89010 scheduling and telemetry features plus the failure label, one row per task
kanvas_phi_curves.csv2,000Learned per-feature curves phi_i(x) of the canonical KAAM, long format
kanvas_benchmark_results.csv6Test metrics for KAAM, MLP and logistic regression

Column definitions are in data_dictionary.md.

Source and construction

Built from the Alibaba Cluster Trace GPU 2020 dataset (Weng et al., "MLaaS in the Wild", NSDI 2022; https://github.com/alibaba/clusterdata). Label: Failed = 1, Terminated = 0. Telemetry features average the previous 1 hour of other workers on the same machine, so no information from the task's own run leaks into the features. User and job identifiers are removed. The raw trace is not redistributed here.

Known limitation

Models trained on this data reach test AUC of about 0.75 with a random split but fall to chance (about 0.45) when the split is grouped by user. One user accounts for 31% of tasks and 54% of failures. Do not use random-split scores from this data as evidence of cross-user generalization.

Citation

Please cite the original dataset paper (Weng et al., NSDI 2022) and link the GitHub repository above.