datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dr.Sparse-GPT56Luna-eval-b200-rl562-spgemm
Dr.Sparse — GPT-5.6 Luna SpGEMM eval on the 562-matrix pool
Snapshot state: complete (generated 2026-09-19 08:03 UTC)
phase
results
explore (3 branches x 5 iterations)
1686 / 1686
exploit (top-2 branches x 10 iterations)
1124 / 1124
matrices covered
562 / 562
What this is
Tree-search SpGEMM kernel optimization over the 562 matrices of
KinGeorge/Dr.Sparse-RL-train-562,
run on NVIDIA B200 (sm_100) and scored against cuSPARSE.
Model:… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesChen122/Dr.Sparse-GPT56Luna-eval-b200-rl562-spgemm.sparseatt-fineweb-large-tokenized
SparseAtt FineWeb Large Tokenized Export
This repository contains the pretokenized training export used by SparseAtt. The source corpus is FineWeb, distributed under ODC-By 1.0; downstream users must also follow the applicable Common Crawl terms.
The full export has 12,748,481 records of 2,048 token IDs each (Qwen3 tokenizer vocabulary size 151,936). The training loader packs each group of four records into an 8,192-token sequence. This repository is being populated… See the full description on the dataset page: https://huggingface.co/datasets/rubentium/sparseatt-fineweb-large-tokenized.Dr.Sparse-SFT-luna10-v4b
Dr.Sparse SFT data — Luna-10 distillation, round v4b (Qwen3.5-9B student)
The exact data behind the qwen3.5-9b-sft-v4b student (OTF-81: 46 % of held-out matrices solved with
reasoning on, 58 % with reasoning off; the base model solves 0). Built from GPT-5.6 ("Luna")
trajectories on the 562-matrix Dr.Sparse training pool; no OTF test-set matrix appears anywhere.
path
what
rows
qwen3.5-9b-v4b/{train,val}.parquet
train on this. Rendered for the Qwen3.5 chat template:… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-SFT-luna10-v4b.stack-v2-sparse-classes-36k
Stack v2 Sparse Python Classes 36k
This is a 36,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments.
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters.
Splits
train.jsonl: 35,000
val.jsonl: 500
test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-36k.stack-v2-sparse-classes-10k
Stack v2 Sparse Python Classes 10k
This is a 10,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments.
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters.
Splits
train.jsonl: 9,000
val.jsonl: 500
test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-10k.stack-v2-sparse-classes-75kplus
Stack v2 Sparse Python Classes 75kplus
This is a frozen snapshot with 75829 samples for Diffusion + Autoregressive hybrid code generation experiments.
Splits
train.jsonl: 74829
val.jsonl: 500
test.jsonl: 500
all.jsonl: 75829
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-75kplus.sparseatt-mask-cache-8k-mixed
SparseAtt 8k Mixed mask cache
This repository contains the stored 36-layer mask cache from SparseAtt. It is a generated training artifact, not a raw text dataset. The two build marker files are included.
The files preserve the source relative layout at sparse/mask_cache_8k_mixed/. inventory.tsv lists every file, its byte size, and its source SHA-256 digest.
