datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dr.Sparse-SFT-luna10-v4b
Dr.Sparse SFT data — Luna-10 distillation, round v4b (Qwen3.5-9B student)
The exact data behind the qwen3.5-9b-sft-v4b student (OTF-81: 46 % of held-out matrices solved with
reasoning on, 58 % with reasoning off; the base model solves 0). Built from GPT-5.6 ("Luna")
trajectories on the 562-matrix Dr.Sparse training pool; no OTF test-set matrix appears anywhere.
path
what
rows
qwen3.5-9b-v4b/{train,val}.parquet
train on this. Rendered for the Qwen3.5 chat template:… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-SFT-luna10-v4b.stack-v2-sparse-classes-36k
Stack v2 Sparse Python Classes 36k
This is a 36,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments.
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters.
Splits
train.jsonl: 35,000
val.jsonl: 500
test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-36k.stack-v2-sparse-classes-10k
Stack v2 Sparse Python Classes 10k
This is a 10,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments.
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters.
Splits
train.jsonl: 9,000
val.jsonl: 500
test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-10k.stack-v2-sparse-classes-75kplus
Stack v2 Sparse Python Classes 75kplus
This is a frozen snapshot with 75829 samples for Diffusion + Autoregressive hybrid code generation experiments.
Splits
train.jsonl: 74829
val.jsonl: 500
test.jsonl: 500
all.jsonl: 75829
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-75kplus.sparseatt-mask-cache-8k-mixed
SparseAtt 8k Mixed mask cache
This repository contains the stored 36-layer mask cache from SparseAtt. It is a generated training artifact, not a raw text dataset. The two build marker files are included.
The files preserve the source relative layout at sparse/mask_cache_8k_mixed/. inventory.tsv lists every file, its byte size, and its source SHA-256 digest.
