Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01charge-benchmark /Charge-040_0040-Sparse-Mono0 likes9.3k downloads1y agoHugging Face02DesmondYMTang2024 /Language-Grounded_Sparse_Encoder_Training Language-Grounded Sparse Encoder (LanSE) — Training Data This repository hosts the AI-generated images and human annotation datasets accompanying the paper: Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.textimage-classification100K<n<1M1 likes5.4k downloads1mo agoHugging Face03DiogenesChen122 /Dr.Sparse-Granite42-8B-eval-b200-otf81-spgemm Dr.Sparse — Granite 4.2 8B SpGEMM baseline (OTF-81) ibm-granite/granite-4.2-8b 在 Dr.Sparse OTF 保留测试集上的 SpGEMM baseline。 levels 1-3(排除 level4),单轨迹无树搜索,每矩阵 10 轮迭代,B200 (sm_100)。 结果 level 矩阵 正确 跑赢 cuSPARSE 中位加速比 最大 level1_small 8 1 0 0.244 0.24 level2_medium 33 3 0 0.024 0.62 level3_large 38 4 0 0.092 1.00 合计 79 8 0 0.102 1.00 79 个矩阵里 8 个产出正确 kernel,无一跑赢 cuSPARSE。中位加速比 0.102 表示比 cuSPARSE 慢约十倍;最好的一个仅持平。 编译失败的错误类型分散:cudaMalloc 重载不匹配、const 限定符未去除、… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesChen122/Dr.Sparse-Granite42-8B-eval-b200-otf81-spgemm.0 likes5k downloads8d agoHugging Face04DiogenesChen122 /Dr.Sparse-Gemma4-12B-eval-b200-otf81-spgemm Dr.Sparse — Gemma 4 12B base, SpGEMM on OTF-81 google/gemma-4-12B-it(未经微调)在 Dr.Sparse OTF 保留测试集上的 SpGEMM 评测。 levels 1-3(排除 level4),单轨迹无树搜索,每矩阵 10 轮迭代,B200 (sm_100)。 结果 level 矩阵 正确 跑赢 cuSPARSE 正确者中位加速比 最大 level1_small 8 0 0 - - level2_medium 31 0 0 - - level3_large 36 1 0 0.057 0.06 合计 75 1 0 0.057 0.06 81 个矩阵中 75 个跑完。其余 6 个因模型上下文超限(vLLM 返回 400)中止,未计入。 失败以真实编译错误为主,例如在 __shared__ 变量上写初始值、函数重复定义、 引用不存在的结构体成员。 对照 同一数据上的 SFT 版本见… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesChen122/Dr.Sparse-Gemma4-12B-eval-b200-otf81-spgemm.0 likes3.5k downloads5d agoHugging Face05DiogenesChen122 /Dr.Sparse-Gemma4-12B-SFT-eval-b200-otf81-spgemm Dr.Sparse — Gemma 4 12B SFT, SpGEMM on OTF-81 Gemma 4 12B 经 Luna-10 v4b 数据 LoRA SFT 后,在 Dr.Sparse OTF 保留测试集上的 SpGEMM 评测。 levels 1-3(排除 level4),单轨迹无树搜索,每矩阵 10 轮迭代,B200 (sm_100)。 模型权重:DiogenesChen122/Dr.Sparse-Gemma4-12B-SFT-luna10-v4b 结果 level 矩阵 正确 跑赢 cuSPARSE 正确者中位加速比 最大 level1_small 8 3 0 0.263 0.49 level2_medium 34 6 2 0.373 2.15 level3_large 37 3 0 0.221 0.88 合计 79 12 2 0.314 2.15 81 个矩阵中 79 个跑完。其余 2 个因模型上下文超限(vLLM 返回 400)中止,未计入。… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesChen122/Dr.Sparse-Gemma4-12B-SFT-eval-b200-otf81-spgemm.0 likes3.4k downloads5d agoHugging Face06OpenDriveLab /SparseVideoNav SparseVideoNav Datasets This repository contains the real-world navigation datasets released with OpenDriveLab/SparseVideoNav: BVN: Beyond-the-View Navigation. IFN: Instruction-Following Navigation. Project links: Project page: https://opendrivelab.com/SparseVideoNav GitHub: https://github.com/OpenDriveLab/SparseVideoNav Paper: https://arxiv.org/abs/2602.05827 Dataset Summary SparseVideoNav studies real-world vision-language navigation with sparse future… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab/SparseVideoNav.tabularrobotics10K<n<100K4 likes3.1k downloads1mo agoHugging Face07P2SAMAPA /p2-etf-sparse-elasticnet-results0 likes1.5k downloads20h agoHugging Face08jack635 /sparse-metric-anchors-ycb Sparse Metric Anchors — YCB benchmark The 42-object benchmark behind the paper Sparse Metric Anchors for a Single-View 3D Generative Prior: The Output Frame Is the Bottleneck (ISIR, Sorbonne Université, 2026). Code and paper: github.com/635jack/sparse-metric-anchors — its colab/reproduce.ipynb recomputes every table of the paper from this dataset on a CPU runtime. The paper asks what limits the injection of a few metric measurements — tactile contacts, one depth map — into a… See the full description on the dataset page: https://huggingface.co/datasets/jack635/sparse-metric-anchors-ycb.imageimage-to-3dn<1K0 likes796 downloads24d agoHugging Face09fin-ai-lab /Market-1T-1Hz-2019H2-2020-sparse Market-1T — 1 Hz, July 2019 → December 2020 (sparse) Code: fin-ai-lab/tfwm downloads this data, trains the TFWM encoders on it, and reproduces the paper's tables. 230,025 ticker-days of 1 Hz US equity market data over 376 trading days (2019-07-01 → 2020-12-31). The window deliberately straddles the February–March 2020 COVID crash and the recovery through 2020, so it contains a genuine regime break rather than a single stationary market. Three layouts, three… See the full description on the dataset page: https://huggingface.co/datasets/fin-ai-lab/Market-1T-1Hz-2019H2-2020-sparse.time-series-forecasting100K<n<1M0 likes551 downloads10d agoHugging Face10CodeMasterCody3D /glm46v-flash-ultramega-sparse-hmap0 likes551 downloads19d agoHugging Face11justicedao /open-us-law-sparse-graphrag State laws snapshot sparse GraphRAG CID-keyed retrieval release in the same thin-client layout as Publicus/skillcenter-ir: Zstandard Parquet shards of at most 4,096 rows entry_cid as the canonical content identity; document_index is a compact pointer compact BM25 term-range, vector-centroid, corpus, and adjacency routing indexes thenlper/gte-small 384-d L2 vectors, cosine-sorted inside centroid shards queries fetch the manifest, routing indexes, and only the routed shards This… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/open-us-law-sparse-graphrag.text-retrieval100K<n<1M0 likes517 downloads12d agoHugging Face12panwh /SparseCam4D1 likes402 downloads6mo agoHugging Face13charge-benchmark /Charge-010_0050-Sparse-Mono0 likes395 downloads1y agoHugging Face14serteal /sparse-probing Sparse Probing Datasets 155 binary classification tasks for probing language model representations. From: "Are Sparse Autoencoders Useful? A Case Study in Sparse Probing" (arXiv:2502.16681) Source: EleutherAI/sae-probes Usage from datasets import load_dataset # Load a specific dataset ds = load_dataset("serteal/sparse-probing", "87_glue_cola") # List available configurations from datasets import get_dataset_config_names configs =… See the full description on the dataset page: https://huggingface.co/datasets/serteal/sparse-probing.text100K<n<1M0 likes363 downloads8mo agoHugging Face15rmems /sparse-reward-long-tasks Sparse Reward Long Tasks Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/sparse-reward-long-tasks.text1K<n<10K0 likes332 downloads17d agoHugging Face16Angshul /SparseGeometricRAG SparseGeometricRAG CPU-first sparse geometric retrieval for practical top-10 RAG No transformer inference at retrieval time. No retrieval GPU requirement. No dense document-vector dot products. No external API. SparseGeometricRAG is a retrieval system built around one systems objective: make the retrieval layer cheap enough to run on ordinary multicore CPU hardware without turning the corpus into a dense embedding database. It uses sparse TF-IDF geometry, fuzzy… See the full description on the dataset page: https://huggingface.co/datasets/Angshul/SparseGeometricRAG.0 likes323 downloads2mo agoHugging Face17iridescentttt /SparseEval_benchmark_data Benchmark Data This directory contains the raw benchmark prediction results in CSV format. These files represent the model outputs and ground truth correctness for various datasets. File Format Each CSV file should contain the following columns: source: The identifier of the model that generated the prediction. item: The identifier of the specific test instance (question/sample). correct: A binary value indicating whether the model's prediction was correct (1) or… See the full description on the dataset page: https://huggingface.co/datasets/iridescentttt/SparseEval_benchmark_data.0 likes297 downloads8mo agoHugging Face18gist-sparse-attention /GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4-data GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4-data This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8-chunk4. Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4 — model trained on this dataset 0 likes288 downloads6mo agoHugging Face19SparseWake /sparsewake SparseWake SparseWake is a synthetic benchmark for sparse temporal hydrodynamic sensing. ICLR 2027 release The expanded release adds controlled multi-source mixtures and common-prior nearest-source tasks, with complete core data banks, reference checkpoints, a small review supplement, and reproduction code with a frozen wake-library input. Download release iclr2027-v1.0rc2 The version page lists the three archives, exact sizes, checksums, extraction instructions… See the full description on the dataset page: https://huggingface.co/datasets/SparseWake/sparsewake.0 likes277 downloads13d agoHugging Face20rubentium /sparse-resultstext0 likes268 downloads2d agoHugging Face21sumith2425 /TRAIN_SPARSE_20 likes267 downloads8mo agoHugging Face22gist-sparse-attention /GSA-PT-Llama-3.2-1B-chunk8-data GSA-PT-Llama-3.2-1B-chunk8-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk8. Each sample is tokenized and formatted with GSA gist tokens for continued pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Models yuzhenm/GSA-PT-Llama-3.2-1B-chunk8 — model trained on this dataset 0 likes252 downloads6mo agoHugging Face23amallia /sparse0 likes242 downloads1y agoHugging Face24maeyounes /SparseCraft-dataset SparseCraft [ECCV'24] SparseCraft: Few-Shot Neural Reconstruction through Stereopsis Guided Geometric Linearization Project DTU Dataset We provide preprocessed DTU data and results for the tasks of novel view synthesis and surface reconstruction. It contains the following directories: sparsecraft_data ├── nvs # Novel View Synthesis task data and results │ └── mvs_data │ ├── scan103 │ ├── ... │ └── results # Results for training using… See the full description on the dataset page: https://huggingface.co/datasets/maeyounes/SparseCraft-dataset.imagen<1K0 likes216 downloads2y agoHugging Face25DiogenesChen122 /Dr.Sparse-Ornith15-9B-eval-b200-otf-spgemm-partial Dr.Sparse — Ornith-1.5-9B SpGEMM baseline (partial, 12/81 matrices) Partial baseline of ornith-ai/Ornith-1.5-9B on the Dr.Sparse OTF held-out test set, SpGEMM only, levels 1-3 (level4 excluded). B200, single trajectory (no tree search). Why this run is partial The run was stopped after 12 of 81 matrices. HiPerGator terminates jobs that hold a GPU without using it, and this eval layout gives each matrix its own GPU while the agent spends most of each iteration… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesChen122/Dr.Sparse-Ornith15-9B-eval-b200-otf-spgemm-partial.0 likes199 downloads18d agoHugging Face26mim-chess-vlas /train_800_sparse__mask__overlay_a75__sim__all_cameras__staticThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "observation.state": { "dtype": "float32", "shape": [ 9 ], "names": [ "x", "y", "z", "qx", "qy", "qz", "qw", "g1", "g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__overlay_a75__sim__all_cameras__static.tabularrobotics100K<n<1M0 likes187 downloads14d agoHugging Face27gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk32-data GSA-PT-Qwen2-7B-Instruct-chunk32-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk32. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk32 — model trained on this dataset 0 likes183 downloads6mo agoHugging Face28mim-chess-vlas /train_800_sparse__no_mask__ur5eThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "observation.state": { "dtype": "float32", "shape": [ 9 ], "names": [ "x", "y", "z", "qx", "qy", "qz", "qw", "g1", "g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__no_mask__ur5e.tabularrobotics100K<n<1M0 likes173 downloads21d agoHugging Face29mim-chess-vlas /train_800_sparse__bbox__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "observation.state": { "dtype": "float32", "shape": [ 9 ], "names": [ "x", "y", "z", "qx", "qy", "qz", "qw", "g1", "g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__bbox__separate_channel__sim__all_cameras__live.tabularrobotics100K<n<1M0 likes169 downloads26d agoHugging Face30gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk16-data GSA-PT-Qwen2-7B-Instruct-chunk16-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk16. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk16 — model trained on this dataset tabular10K<n<100K0 likes165 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.