datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datasets-tests-compressioncodecpilot-compression-decision-dataset
CodecPilot Compression-Decision Dataset
用于同格式图片压缩参数预测的原图和编码候选实测结果。给定图像,在所选质量门槛达标的候选中选择体积较小的参数。
This dataset contains input images and measured same-format compression candidates. It contains no trained models and does not store every candidate's encoded output image.
当前正式标注:五个完整数据库
布局版本 v0.7,完整合并发布日期 2026-10-01。原标注与正式增量已按格式合并。
格式
图片文件数
候选组数/张
实测结果数
状态
PNG
8,832
42
370,944
完整
JPEG
18,742
48
899,616
完整
WebP
15,000
45
675,000
完整
JXL
15,000
47
705,000… See the full description on the dataset page: https://huggingface.co/datasets/winrisef/codecpilot-compression-decision-dataset.LLM_compression_calibration
LLM Compression Calibration dataset
This dataset is the default calibration dataset used by Neural Magic for one-shot compression of Large Language Models (LLMs).
Note: This dataset is the result of active research and subject to change without notice.
Dataset Details
Dataset Sources
The current version of this dataset is compiled from data from these datasets:
garage-bAInd/Open-Platypus: 10,000 samples
Data Fields
The dataset contains 2 data… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/LLM_compression_calibration.lean-proof-compression
LeanPolish: Verified Supervision for Lean Proof Compression
A dataset of Lean 4 proof rewrite pairs produced by LeanPolish,
a kernel-verified proof-shortening tool. Every accepted
(original, replacement) pair was kernel-checked under Lean 4.21.0
with Mathlib v4.21.0 before emission, and the rewritten file was
re-elaborated end-to-end by a separate out-of-process verifier.
The dataset is suitable for training models that learn to compress,
simplify, or select proof tactics, and… See the full description on the dataset page: https://huggingface.co/datasets/leanpolish-anon/lean-proof-compression.comma_video_compression_challenge_pr_archive
comma video compression challenge - PR archive corpus
Card last refreshed: 2026-05-11 (companion research artifacts section
added).
This dataset captures every scored Pull Request submitted to
commaai/comma_video_compression_challenge,
the public 2026 contest to compress comma's 0.mkv reference dashcam video
under perceptual + temporal scorer constraints.
For each scored PR we publish:
archive.zip - the exact compressed-archive bytes that were scored
by the contest evaluation… See the full description on the dataset page: https://huggingface.co/datasets/adpena/comma_video_compression_challenge_pr_archive.l1-exact-qwen3-1.7b-compression-1k-8k-alpha3e-4-bs32-n8-32k-t1-146102-rollouts
RL training rollouts
l1_exact_Qwen3-1.7B_compression_n1000to8000_alpha3e-4_bs32_n8_32k_1epoch
One verified gzip JSONL shard per training step; 256 responses per shard.
L1-Exact: MathVerify correctness after thinking without an EOS gate, plus the original linear target-length reward. Fixed random targets 1000..8000; alpha=0.0003.
l0-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-seqmean-verl091-146102-rollouts
RL training rollouts
l0_Qwen3-1.7B_compression_bs32_n16_32k_t1_no_eos_verl091_seqmean
One verified gzip JSONL shard per training step; 512 responses per shard.
Historical compression MathVerify after thinking, without an EOS gate.
f-cov-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-verl091-146103-rollouts
RL training rollouts
f_cov_l0_0_no_eos_Qwen3-1.7B_compression_bs32_n16_32k_t1_verl091
One verified gzip JSONL shard per training step; 512 responses per shard.
Historical compression MathVerify after thinking, without an EOS gate.
l1-exact-qwen3-1.7b-compression-bs32-n8-32k-t1-146103-rollouts
RL training rollouts
l1_exact_Qwen3-1.7B_compression_n400to12000_bs32_n8_32k_1epoch
One verified gzip JSONL shard per training step; 256 responses per shard.
L1-Exact: MathVerify correctness after thinking without an EOS gate, plus the original linear target-length reward. Fixed targets 400..12000; alpha=0.000075.
ImageNet-C-jpeg_compression-severity_5l4096-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-verl091-146102-rollouts
RL training rollouts
l4096_Qwen3-1.7B_compression_bs32_n16_32k_t1_no_eos_verl091
One verified gzip JSONL shard per training step; 512 responses per shard.
Historical compression MathVerify after thinking, without an EOS gate.
l1-exact-qwen3-1.7b-compression-100-4k-alpha3e-4-bs32-n8-32k-t1-146103-rollouts
RL training rollouts
l1_exact_Qwen3-1.7B_compression_n100to4000_alpha3e-4_bs32_n8_32k_1epoch
One verified gzip JSONL shard per training step; 256 responses per shard.
L1-Exact: MathVerify correctness after thinking without an EOS gate, plus the original linear target-length reward. Fixed random targets 100..4000; alpha=0.0003.
l1-exact-qwen3-1.7b-compression-100-6k-alpha3e-4-bs32-n8-32k-t1-146102-rollouts
RL training rollouts
l1_exact_Qwen3-1.7B_compression_n100to6000_alpha3e-4_bs32_n8_32k_1epoch
One verified gzip JSONL shard per training step; 256 responses per shard.
L1-Exact: MathVerify correctness after thinking without an EOS gate, plus the original linear target-length reward. Fixed random targets 100..6000; alpha=0.0003.
compression-math-100plus-topup128k-qwen3-1.7b-verl091-146103-20261004
Compression math responses: every question has at least 160k output tokens
585,003 complete responses from five Qwen3-1.7B models on six datasets. Every model-question pair now has at least 163,840 total generated output tokens (160 × 1,024).
This update adds 21,973 complete responses (21,972,891 output tokens) to the previous 563,030-response release. The repository name records the original 128k release; the current data includes the completed 160k supplementation.
Metrics… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/compression-math-100plus-topup128k-qwen3-1.7b-verl091-146103-20261004.f-cov-l4096-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-seqmean-verl091-146102-rollouts
RL training rollouts
f_cov_l0_4096_no_eos_Qwen3-1.7B_compression_bs32_n16_32k_t1_seqmean_verl091
One verified gzip JSONL shard per training step; 512 responses per shard.
Historical compression MathVerify after thinking, without an EOS gate.
fastwam-lerobotl0-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-verl091-146102-rollouts
RL training rollouts
l0_Qwen3-1.7B_compression_bs32_n16_32k_t1_no_eos_verl091
One verified gzip JSONL shard per training step; 512 responses per shard.
Historical compression MathVerify after thinking, without an EOS gate.
sentence-compression
Dataset Card for "sentence-compression"
Dataset Summary
Dataset with pairs of equivalent sentences.
The dataset is provided "AS IS" without any warranty, express or implied.
Google disclaims all liability for any damages, direct or indirect, resulting from using the dataset.
Disclaimer: The team releasing sentence-compression did not upload the dataset to the Hub and did not write a dataset card. These steps were done by the Hugging Face team.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/sentence-compression.msr_text_compressionThis dataset contains sentences and short paragraphs with corresponding shorter (compressed) versions. There are up to five compressions for each input text, together with quality judgements of their meaning preservation and grammaticality. The dataset is derived using source texts from the Open American National Corpus (ww.anc.org) and crowd-sourcing.vision-token-compression-bench
OPTIC-Bench
Optical Text In-Context Benchmark: how reliably do LLMs consume text
delivered as rendered images versus plain text tokens?
In summary, the evaluation reported here finds that optical text compression
is effective only within a narrow and specific envelope. Delivering content
as rendered images genuinely reduces input tokens, by thirteen to
fifty-four per cent depending on the model and the language, but only when
the document is long, the rendering is dense and the… See the full description on the dataset page: https://huggingface.co/datasets/translorentz/vision-token-compression-bench.compression-pretraining-data
Dataset
Each example contains prompt (chat format) and target fields.
from datasets import load_dataset
ds = load_dataset("leonli66/compression-pretraining-data", "<config_name>")
svs-lame-compression-jpeg-vs-neural
Comparaison visuelle : compression neuronale vs JPEG sur lames histopathologiques
Ce dataset permet à un anatomopathologiste de juger à l'œil nu si une image
de lame numérique compressée par un réseau de neurones est visuellement
équivalente à la même lame compressée en JPEG (qui est le standard)
En une phrase
On a pris 5 lames histopathologiques au format SVS, on les a compressées avec
4 modèles neuronaux et avec JPEG Q75, à
deux niveaux d'agressivité (q5 ≈… See the full description on the dataset page: https://huggingface.co/datasets/nathbns/svs-lame-compression-jpeg-vs-neural.compression-eval-datasets
Compression Evaluation Datasets
Common test sets for learned image compression evaluation, packaged for easy download.
Contents
Folder
Description
Images
Size
kodak/
Kodak PhotoCD (kodim01–kodim24)
24
~15MB
tecnick/
Tecnick RGB test images (1200×1200)
40
~66MB
clic2021_valid/
CLIC professional validation images (local folder name: CLIC2021_valid)
41
~129MB
Download
# Full dataset
hf download SCU-VIP-Lab/compression-eval-datasets… See the full description on the dataset page: https://huggingface.co/datasets/SCU-VIP-Lab/compression-eval-datasets.round-trip-code-compressioncorruption-jpeg_compression
Corruption Dataset: Jpeg_Compression
Dataset Description
This dataset contains corrupted versions of ImageNet-1K images using jpeg_compression corruption. It is part of the ImageNet-C benchmark for evaluating model robustness to common image corruptions.
Dataset Structure
Train: 1,281,167 corrupted images
Validation: 50,000 corrupted images
Classes: 1000 ImageNet-1K classes
Format: Arrow (Hugging Face Datasets)
Corruption Type: Jpeg_Compression… See the full description on the dataset page: https://huggingface.co/datasets/MarMaster/corruption-jpeg_compression.lingbot-va-attn
LingBot-VA Attention Analysis Dataset
Attention-analysis dataset generated on h100-server for the RoboTwin task grab-the-medium-sized-white-mug-rotate-it-place-it-on-the-table-and-hook-it-onto-the-smooth-dark-gray-rack, using the LingBot-VA 1.0-style velocity FlowMatch inference path (checkpoint lingbot-va-posttrain-robotwin).
Content
artifacts/archives/lingbot-va-attn-trajectory-6steps.tar.gz.part00 … part09 — the full six-step raw attention capture: 4,320 dense… See the full description on the dataset page: https://huggingface.co/datasets/kv-compression/lingbot-va-attn.all-deletion-compressionswikipedia-deletion-compressionsfw-edu-cl100kParameter-Golf-V17-512Cube-XYZ-Priority-Orbital-Compression
Parameter Golf V17 — 512Cube Solution Bank
This is an English research-control dataset for OpenAI Parameter Golf work. It is not a replacement for FineWeb and must not be used as a substitute training or validation corpus. FineWeb remains the canonical data path for contest scoring.
The dataset captures three things:
V17 512Cube routing concepts translated into English.
Contest and submission guardrails for legal, reproducible BPB reduction.
Screenshot-derived scouting observations… See the full description on the dataset page: https://huggingface.co/datasets/8Planetterraforming/Parameter-Golf-V17-512Cube-XYZ-Priority-Orbital-Compression.
