compression
datasets-tests-compressioncodecpilot-compression-decision-dataset
CodecPilot Compression-Decision Dataset
用于同格式图片压缩参数预测的原图和编码候选实测结果。给定图像,在所选质量门槛达标的候选中选择体积较小的参数。
This dataset contains input images and measured same-format compression candidates. It contains no trained models and does not store every candidate's encoded output image.
当前正式标注:五个完整数据库
布局版本 v0.7,完整合并发布日期 2026-10-01。原标注与正式增量已按格式合并。
格式
图片文件数
候选组数/张
实测结果数
状态
PNG
8,832
42
370,944
完整
JPEG
18,742
48
899,616
完整
WebP
15,000
45
675,000
完整
JXL
15,000
47
705,000… See the full description on the dataset page: https://huggingface.co/datasets/winrisef/codecpilot-compression-decision-dataset.LLM_compression_calibration
LLM Compression Calibration dataset
This dataset is the default calibration dataset used by Neural Magic for one-shot compression of Large Language Models (LLMs).
Note: This dataset is the result of active research and subject to change without notice.
Dataset Details
Dataset Sources
The current version of this dataset is compiled from data from these datasets:
garage-bAInd/Open-Platypus: 10,000 samples
Data Fields
The dataset contains 2 data… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/LLM_compression_calibration.lean-proof-compression
LeanPolish: Verified Supervision for Lean Proof Compression
A dataset of Lean 4 proof rewrite pairs produced by LeanPolish,
a kernel-verified proof-shortening tool. Every accepted
(original, replacement) pair was kernel-checked under Lean 4.21.0
with Mathlib v4.21.0 before emission, and the rewritten file was
re-elaborated end-to-end by a separate out-of-process verifier.
The dataset is suitable for training models that learn to compress,
simplify, or select proof tactics, and… See the full description on the dataset page: https://huggingface.co/datasets/leanpolish-anon/lean-proof-compression.comma_video_compression_challenge_pr_archive
comma video compression challenge - PR archive corpus
Card last refreshed: 2026-05-11 (companion research artifacts section
added).
This dataset captures every scored Pull Request submitted to
commaai/comma_video_compression_challenge,
the public 2026 contest to compress comma's 0.mkv reference dashcam video
under perceptual + temporal scorer constraints.
For each scored PR we publish:
archive.zip - the exact compressed-archive bytes that were scored
by the contest evaluation… See the full description on the dataset page: https://huggingface.co/datasets/adpena/comma_video_compression_challenge_pr_archive.l1-exact-qwen3-1.7b-compression-1k-8k-alpha3e-4-bs32-n8-32k-t1-146102-rollouts
RL training rollouts
l1_exact_Qwen3-1.7B_compression_n1000to8000_alpha3e-4_bs32_n8_32k_1epoch
One verified gzip JSONL shard per training step; 256 responses per shard.
L1-Exact: MathVerify correctness after thinking without an EOS gate, plus the original linear target-length reward. Fixed random targets 1000..8000; alpha=0.0003.
