datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Charge-040_0040-Sparse-MonoLanguage-Grounded_Sparse_Encoder_Training
Language-Grounded Sparse Encoder (LanSE) — Training Data
This repository hosts the AI-generated images and human annotation datasets accompanying the paper:
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu
National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.Dr.Sparse-Granite42-8B-eval-b200-otf81-spgemm
Dr.Sparse — Granite 4.2 8B SpGEMM baseline (OTF-81)
ibm-granite/granite-4.2-8b 在 Dr.Sparse OTF 保留测试集上的 SpGEMM baseline。
levels 1-3(排除 level4),单轨迹无树搜索,每矩阵 10 轮迭代,B200 (sm_100)。
结果
level
矩阵
正确
跑赢 cuSPARSE
中位加速比
最大
level1_small
8
1
0
0.244
0.24
level2_medium
33
3
0
0.024
0.62
level3_large
38
4
0
0.092
1.00
合计
79
8
0
0.102
1.00
79 个矩阵里 8 个产出正确 kernel,无一跑赢 cuSPARSE。中位加速比 0.102
表示比 cuSPARSE 慢约十倍;最好的一个仅持平。
编译失败的错误类型分散:cudaMalloc 重载不匹配、const 限定符未去除、… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesChen122/Dr.Sparse-Granite42-8B-eval-b200-otf81-spgemm.Dr.Sparse-Gemma4-12B-eval-b200-otf81-spgemm
Dr.Sparse — Gemma 4 12B base, SpGEMM on OTF-81
google/gemma-4-12B-it(未经微调)在 Dr.Sparse OTF 保留测试集上的 SpGEMM 评测。
levels 1-3(排除 level4),单轨迹无树搜索,每矩阵 10 轮迭代,B200 (sm_100)。
结果
level
矩阵
正确
跑赢 cuSPARSE
正确者中位加速比
最大
level1_small
8
0
0
-
-
level2_medium
31
0
0
-
-
level3_large
36
1
0
0.057
0.06
合计
75
1
0
0.057
0.06
81 个矩阵中 75 个跑完。其余 6 个因模型上下文超限(vLLM 返回 400)中止,未计入。
失败以真实编译错误为主,例如在 __shared__ 变量上写初始值、函数重复定义、
引用不存在的结构体成员。
对照
同一数据上的 SFT 版本见… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesChen122/Dr.Sparse-Gemma4-12B-eval-b200-otf81-spgemm.Dr.Sparse-Gemma4-12B-SFT-eval-b200-otf81-spgemm
Dr.Sparse — Gemma 4 12B SFT, SpGEMM on OTF-81
Gemma 4 12B 经 Luna-10 v4b 数据 LoRA SFT 后,在 Dr.Sparse OTF 保留测试集上的 SpGEMM 评测。
levels 1-3(排除 level4),单轨迹无树搜索,每矩阵 10 轮迭代,B200 (sm_100)。
模型权重:DiogenesChen122/Dr.Sparse-Gemma4-12B-SFT-luna10-v4b
结果
level
矩阵
正确
跑赢 cuSPARSE
正确者中位加速比
最大
level1_small
8
3
0
0.263
0.49
level2_medium
34
6
2
0.373
2.15
level3_large
37
3
0
0.221
0.88
合计
79
12
2
0.314
2.15
81 个矩阵中 79 个跑完。其余 2 个因模型上下文超限(vLLM 返回 400)中止,未计入。… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesChen122/Dr.Sparse-Gemma4-12B-SFT-eval-b200-otf81-spgemm.SparseVideoNav
SparseVideoNav Datasets
This repository contains the real-world navigation datasets released with OpenDriveLab/SparseVideoNav:
BVN: Beyond-the-View Navigation.
IFN: Instruction-Following Navigation.
Project links:
Project page: https://opendrivelab.com/SparseVideoNav
GitHub: https://github.com/OpenDriveLab/SparseVideoNav
Paper: https://arxiv.org/abs/2602.05827
Dataset Summary
SparseVideoNav studies real-world vision-language navigation with sparse future… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab/SparseVideoNav.p2-etf-sparse-elasticnet-resultssparse-metric-anchors-ycb
Sparse Metric Anchors — YCB benchmark
The 42-object benchmark behind the paper Sparse Metric Anchors for a Single-View 3D
Generative Prior: The Output Frame Is the Bottleneck (ISIR, Sorbonne Université,
2026). Code and paper: github.com/635jack/sparse-metric-anchors — its colab/reproduce.ipynb recomputes every table of the paper from this dataset on a CPU runtime.
The paper asks what limits the injection of a few metric measurements — tactile
contacts, one depth map — into a… See the full description on the dataset page: https://huggingface.co/datasets/jack635/sparse-metric-anchors-ycb.Market-1T-1Hz-2019H2-2020-sparse
Market-1T — 1 Hz, July 2019 → December 2020 (sparse)
Code: fin-ai-lab/tfwm downloads this data, trains the TFWM
encoders on it, and reproduces the paper's tables.
230,025 ticker-days of 1 Hz US equity market data over 376 trading days
(2019-07-01 → 2020-12-31). The window deliberately straddles the February–March
2020 COVID crash and the recovery through 2020, so it contains a genuine regime
break rather than a single stationary market.
Three layouts, three… See the full description on the dataset page: https://huggingface.co/datasets/fin-ai-lab/Market-1T-1Hz-2019H2-2020-sparse.glm46v-flash-ultramega-sparse-hmapopen-us-law-sparse-graphrag
State laws snapshot sparse GraphRAG
CID-keyed retrieval release in the same thin-client layout as
Publicus/skillcenter-ir:
Zstandard Parquet shards of at most 4,096 rows
entry_cid as the canonical content identity; document_index is a compact pointer
compact BM25 term-range, vector-centroid, corpus, and adjacency routing indexes
thenlper/gte-small 384-d L2 vectors, cosine-sorted inside centroid shards
queries fetch the manifest, routing indexes, and only the routed shards
This… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/open-us-law-sparse-graphrag.SparseCam4DCharge-010_0050-Sparse-Monosparse-probing
Sparse Probing Datasets
155 binary classification tasks for probing language model representations.
From: "Are Sparse Autoencoders Useful? A Case Study in Sparse Probing" (arXiv:2502.16681)
Source: EleutherAI/sae-probes
Usage
from datasets import load_dataset
# Load a specific dataset
ds = load_dataset("serteal/sparse-probing", "87_glue_cola")
# List available configurations
from datasets import get_dataset_config_names
configs =… See the full description on the dataset page: https://huggingface.co/datasets/serteal/sparse-probing.sparse-reward-long-tasks
Sparse Reward Long Tasks
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/sparse-reward-long-tasks.SparseGeometricRAG
SparseGeometricRAG
CPU-first sparse geometric retrieval for practical top-10 RAG
No transformer inference at retrieval time. No retrieval GPU requirement. No dense document-vector dot products. No external API.
SparseGeometricRAG is a retrieval system built around one systems objective: make the retrieval layer cheap enough to run on ordinary multicore CPU hardware without turning the corpus into a dense embedding database. It uses sparse TF-IDF geometry, fuzzy… See the full description on the dataset page: https://huggingface.co/datasets/Angshul/SparseGeometricRAG.SparseEval_benchmark_data
Benchmark Data
This directory contains the raw benchmark prediction results in CSV format. These files represent the model outputs and ground truth correctness for various datasets.
File Format
Each CSV file should contain the following columns:
source: The identifier of the model that generated the prediction.
item: The identifier of the specific test instance (question/sample).
correct: A binary value indicating whether the model's prediction was correct (1) or… See the full description on the dataset page: https://huggingface.co/datasets/iridescentttt/SparseEval_benchmark_data.GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4-data
GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4-data
This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8-chunk4.
Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4 — model trained on this dataset
sparsewake
SparseWake
SparseWake is a synthetic benchmark for sparse temporal hydrodynamic sensing.
ICLR 2027 release
The expanded release adds controlled multi-source mixtures and common-prior nearest-source tasks, with complete core data banks, reference checkpoints, a small review supplement, and reproduction code with a frozen wake-library input.
Download release iclr2027-v1.0rc2
The version page lists the three archives, exact sizes, checksums, extraction instructions… See the full description on the dataset page: https://huggingface.co/datasets/SparseWake/sparsewake.sparse-resultsTRAIN_SPARSE_2GSA-PT-Llama-3.2-1B-chunk8-data
GSA-PT-Llama-3.2-1B-chunk8-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk8.
Each sample is tokenized and formatted with GSA gist tokens for continued pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Models
yuzhenm/GSA-PT-Llama-3.2-1B-chunk8 — model trained on this dataset
sparseSparseCraft-dataset
SparseCraft
[ECCV'24] SparseCraft: Few-Shot Neural Reconstruction through Stereopsis Guided Geometric Linearization
Project
DTU Dataset
We provide preprocessed DTU data and results for the tasks of novel view synthesis and surface reconstruction.
It contains the following directories:
sparsecraft_data
├── nvs # Novel View Synthesis task data and results
│ └── mvs_data
│ ├── scan103
│ ├── ...
│ └── results # Results for training using… See the full description on the dataset page: https://huggingface.co/datasets/maeyounes/SparseCraft-dataset.Dr.Sparse-Ornith15-9B-eval-b200-otf-spgemm-partial
Dr.Sparse — Ornith-1.5-9B SpGEMM baseline (partial, 12/81 matrices)
Partial baseline of ornith-ai/Ornith-1.5-9B on the Dr.Sparse OTF held-out test set,
SpGEMM only, levels 1-3 (level4 excluded). B200, single trajectory (no tree search).
Why this run is partial
The run was stopped after 12 of 81 matrices. HiPerGator terminates jobs that hold a GPU
without using it, and this eval layout gives each matrix its own GPU while the agent spends
most of each iteration… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesChen122/Dr.Sparse-Ornith15-9B-eval-b200-otf-spgemm-partial.train_800_sparse__mask__overlay_a75__sim__all_cameras__staticThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__overlay_a75__sim__all_cameras__static.GSA-PT-Qwen2-7B-Instruct-chunk32-data
GSA-PT-Qwen2-7B-Instruct-chunk32-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk32.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk32 — model trained on this dataset
train_800_sparse__no_mask__ur5eThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__no_mask__ur5e.train_800_sparse__bbox__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__bbox__separate_channel__sim__all_cameras__live.GSA-PT-Qwen2-7B-Instruct-chunk16-data
GSA-PT-Qwen2-7B-Instruct-chunk16-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk16.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk16 — model trained on this dataset
