Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01albertvillanova /datasets-tests-compressiontextn<1K0 likes43k downloads5y agoHugging Face02winrisef /codecpilot-compression-decision-datasetgated CodecPilot Compression-Decision Dataset 用于同格式图片压缩参数预测的原图和编码候选实测结果。给定图像,在所选质量门槛达标的候选中选择体积较小的参数。 This dataset contains input images and measured same-format compression candidates. It contains no trained models and does not store every candidate's encoded output image. 当前正式标注:五个完整数据库 布局版本 v0.7,完整合并发布日期 2026-10-01。原标注与正式增量已按格式合并。 格式 图片文件数 候选组数/张 实测结果数 状态 PNG 8,832 42 370,944 完整 JPEG 18,742 48 899,616 完整 WebP 15,000 45 675,000 完整 JXL 15,000 47 705,000… See the full description on the dataset page: https://huggingface.co/datasets/winrisef/codecpilot-compression-decision-dataset.10K<n<100K0 likes838 downloads9d agoHugging Face03neuralmagic /LLM_compression_calibration LLM Compression Calibration dataset This dataset is the default calibration dataset used by Neural Magic for one-shot compression of Large Language Models (LLMs). Note: This dataset is the result of active research and subject to change without notice. Dataset Details Dataset Sources The current version of this dataset is compiled from data from these datasets: garage-bAInd/Open-Platypus: 10,000 samples Data Fields The dataset contains 2 data… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/LLM_compression_calibration.text10K<n<100K17 likes694 downloads2y agoHugging Face04leanpolish-anon /lean-proof-compression LeanPolish: Verified Supervision for Lean Proof Compression A dataset of Lean 4 proof rewrite pairs produced by LeanPolish, a kernel-verified proof-shortening tool. Every accepted (original, replacement) pair was kernel-checked under Lean 4.21.0 with Mathlib v4.21.0 before emission, and the rewritten file was re-elaborated end-to-end by a separate out-of-process verifier. The dataset is suitable for training models that learn to compress, simplify, or select proof tactics, and… See the full description on the dataset page: https://huggingface.co/datasets/leanpolish-anon/lean-proof-compression.tabulartext-generation10K<n<100K1 likes554 downloads15d agoHugging Face05adpena /comma_video_compression_challenge_pr_archive comma video compression challenge - PR archive corpus Card last refreshed: 2026-05-11 (companion research artifacts section added). This dataset captures every scored Pull Request submitted to commaai/comma_video_compression_challenge, the public 2026 contest to compress comma's 0.mkv reference dashcam video under perceptual + temporal scorer constraints. For each scored PR we publish: archive.zip - the exact compressed-archive bytes that were scored by the contest evaluation… See the full description on the dataset page: https://huggingface.co/datasets/adpena/comma_video_compression_challenge_pr_archive.image-to-imagen<1K0 likes519 downloads5mo agoHugging Face06hi-todayis-jh /l1-exact-qwen3-1.7b-compression-1k-8k-alpha3e-4-bs32-n8-32k-t1-146102-rollouts RL training rollouts l1_exact_Qwen3-1.7B_compression_n1000to8000_alpha3e-4_bs32_n8_32k_1epoch One verified gzip JSONL shard per training step; 256 responses per shard. L1-Exact: MathVerify correctness after thinking without an EOS gate, plus the original linear target-length reward. Fixed random targets 1000..8000; alpha=0.0003. tabular10K<n<100K0 likes480 downloads6d agoHugging Face07hi-todayis-jh /l0-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-seqmean-verl091-146102-rollouts RL training rollouts l0_Qwen3-1.7B_compression_bs32_n16_32k_t1_no_eos_verl091_seqmean One verified gzip JSONL shard per training step; 512 responses per shard. Historical compression MathVerify after thinking, without an EOS gate. tabular10K<n<100K0 likes461 downloads9d agoHugging Face08hi-todayis-jh /f-cov-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-verl091-146103-rollouts RL training rollouts f_cov_l0_0_no_eos_Qwen3-1.7B_compression_bs32_n16_32k_t1_verl091 One verified gzip JSONL shard per training step; 512 responses per shard. Historical compression MathVerify after thinking, without an EOS gate. tabular10K<n<100K0 likes455 downloads9d agoHugging Face09hi-todayis-jh /l1-exact-qwen3-1.7b-compression-bs32-n8-32k-t1-146103-rollouts RL training rollouts l1_exact_Qwen3-1.7B_compression_n400to12000_bs32_n8_32k_1epoch One verified gzip JSONL shard per training step; 256 responses per shard. L1-Exact: MathVerify correctness after thinking without an EOS gate, plus the original linear target-length reward. Fixed targets 400..12000; alpha=0.000075. tabular10K<n<100K0 likes452 downloads8d agoHugging Face10Tsomaros /ImageNet-C-jpeg_compression-severity_5image10K<n<100K0 likes445 downloads2y agoHugging Face11hi-todayis-jh /l4096-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-verl091-146102-rollouts RL training rollouts l4096_Qwen3-1.7B_compression_bs32_n16_32k_t1_no_eos_verl091 One verified gzip JSONL shard per training step; 512 responses per shard. Historical compression MathVerify after thinking, without an EOS gate. tabular10K<n<100K0 likes439 downloads9d agoHugging Face12hi-todayis-jh /l1-exact-qwen3-1.7b-compression-100-4k-alpha3e-4-bs32-n8-32k-t1-146103-rollouts RL training rollouts l1_exact_Qwen3-1.7B_compression_n100to4000_alpha3e-4_bs32_n8_32k_1epoch One verified gzip JSONL shard per training step; 256 responses per shard. L1-Exact: MathVerify correctness after thinking without an EOS gate, plus the original linear target-length reward. Fixed random targets 100..4000; alpha=0.0003. tabular10K<n<100K0 likes430 downloads5d agoHugging Face13hi-todayis-jh /l1-exact-qwen3-1.7b-compression-100-6k-alpha3e-4-bs32-n8-32k-t1-146102-rollouts RL training rollouts l1_exact_Qwen3-1.7B_compression_n100to6000_alpha3e-4_bs32_n8_32k_1epoch One verified gzip JSONL shard per training step; 256 responses per shard. L1-Exact: MathVerify correctness after thinking without an EOS gate, plus the original linear target-length reward. Fixed random targets 100..6000; alpha=0.0003. tabular10K<n<100K0 likes394 downloads5d agoHugging Face14hi-todayis-jh /compression-math-100plus-topup128k-qwen3-1.7b-verl091-146103-20261004 Compression math responses: every question has at least 160k output tokens 585,003 complete responses from five Qwen3-1.7B models on six datasets. Every model-question pair now has at least 163,840 total generated output tokens (160 × 1,024). This update adds 21,973 complete responses (21,972,891 output tokens) to the previous 563,030-response release. The repository name records the original 128k release; the current data includes the completed 160k supplementation. Metrics… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/compression-math-100plus-topup128k-qwen3-1.7b-verl091-146103-20261004.tabulartext-generation100K<n<1M0 likes392 downloads6d agoHugging Face15hi-todayis-jh /f-cov-l4096-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-seqmean-verl091-146102-rollouts RL training rollouts f_cov_l0_4096_no_eos_Qwen3-1.7B_compression_bs32_n16_32k_t1_seqmean_verl091 One verified gzip JSONL shard per training step; 512 responses per shard. Historical compression MathVerify after thinking, without an EOS gate. tabular10K<n<100K0 likes353 downloads8d agoHugging Face16kv-compression /fastwam-lerobotimage1K<n<10K0 likes313 downloads2mo agoHugging Face17hi-todayis-jh /l0-qwen3-1.7b-compression-bs32-n16-32k-t1-no-eos-verl091-146102-rollouts RL training rollouts l0_Qwen3-1.7B_compression_bs32_n16_32k_t1_no_eos_verl091 One verified gzip JSONL shard per training step; 512 responses per shard. Historical compression MathVerify after thinking, without an EOS gate. tabular10K<n<100K0 likes281 downloads10d agoHugging Face18embedding-data /sentence-compression Dataset Card for "sentence-compression" Dataset Summary Dataset with pairs of equivalent sentences. The dataset is provided "AS IS" without any warranty, express or implied. Google disclaims all liability for any damages, direct or indirect, resulting from using the dataset. Disclaimer: The team releasing sentence-compression did not upload the dataset to the Hub and did not write a dataset card. These steps were done by the Hugging Face team. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/sentence-compression.textsentence-similarity100K<n<1M22 likes231 downloads4y agoHugging Face19microsoft /msr_text_compressionThis dataset contains sentences and short paragraphs with corresponding shorter (compressed) versions. There are up to five compressions for each input text, together with quality judgements of their meaning preservation and grammaticality. The dataset is derived using source texts from the Open American National Corpus (ww.anc.org) and crowd-sourcing.summarization1K<n<10K10 likes222 downloads3y agoHugging Face20translorentz /vision-token-compression-bench OPTIC-Bench Optical Text In-Context Benchmark: how reliably do LLMs consume text delivered as rendered images versus plain text tokens? In summary, the evaluation reported here finds that optical text compression is effective only within a narrow and specific envelope. Delivering content as rendered images genuinely reduces input tokens, by thirteen to fifty-four per cent depending on the model and the language, but only when the document is long, the rendering is dense and the… See the full description on the dataset page: https://huggingface.co/datasets/translorentz/vision-token-compression-bench.imagevisual-question-answering1K<n<10K0 likes216 downloads3mo agoHugging Face21leonli66 /compression-pretraining-data Dataset Each example contains prompt (chat format) and target fields. from datasets import load_dataset ds = load_dataset("leonli66/compression-pretraining-data", "<config_name>") text100M<n<1B0 likes179 downloads10mo agoHugging Face22nathbns /svs-lame-compression-jpeg-vs-neural Comparaison visuelle : compression neuronale vs JPEG sur lames histopathologiques Ce dataset permet à un anatomopathologiste de juger à l'œil nu si une image de lame numérique compressée par un réseau de neurones est visuellement équivalente à la même lame compressée en JPEG (qui est le standard) En une phrase On a pris 5 lames histopathologiques au format SVS, on les a compressées avec 4 modèles neuronaux et avec JPEG Q75, à deux niveaux d'agressivité (q5 ≈… See the full description on the dataset page: https://huggingface.co/datasets/nathbns/svs-lame-compression-jpeg-vs-neural.imagen<1K0 likes168 downloads3mo agoHugging Face23SCU-VIP-Lab /compression-eval-datasets Compression Evaluation Datasets Common test sets for learned image compression evaluation, packaged for easy download. Contents Folder Description Images Size kodak/ Kodak PhotoCD (kodim01–kodim24) 24 ~15MB tecnick/ Tecnick RGB test images (1200×1200) 40 ~66MB clic2021_valid/ CLIC professional validation images (local folder name: CLIC2021_valid) 41 ~129MB Download # Full dataset hf download SCU-VIP-Lab/compression-eval-datasets… See the full description on the dataset page: https://huggingface.co/datasets/SCU-VIP-Lab/compression-eval-datasets.imageothern<1K0 likes164 downloads1mo agoHugging Face24nickil-shay-antonio /round-trip-code-compressiontext100K<n<1M0 likes157 downloads1y agoHugging Face25MarMaster /corruption-jpeg_compression Corruption Dataset: Jpeg_Compression Dataset Description This dataset contains corrupted versions of ImageNet-1K images using jpeg_compression corruption. It is part of the ImageNet-C benchmark for evaluating model robustness to common image corruptions. Dataset Structure Train: 1,281,167 corrupted images Validation: 50,000 corrupted images Classes: 1000 ImageNet-1K classes Format: Arrow (Hugging Face Datasets) Corruption Type: Jpeg_Compression… See the full description on the dataset page: https://huggingface.co/datasets/MarMaster/corruption-jpeg_compression.image-classification1M<n<10M0 likes153 downloads1y agoHugging Face26kv-compression /lingbot-va-attn LingBot-VA Attention Analysis Dataset Attention-analysis dataset generated on h100-server for the RoboTwin task grab-the-medium-sized-white-mug-rotate-it-place-it-on-the-table-and-hook-it-onto-the-smooth-dark-gray-rack, using the LingBot-VA 1.0-style velocity FlowMatch inference path (checkpoint lingbot-va-posttrain-robotwin). Content artifacts/archives/lingbot-va-attn-trajectory-6steps.tar.gz.part00 … part09 — the full six-step raw attention capture: 4,320 dense… See the full description on the dataset page: https://huggingface.co/datasets/kv-compression/lingbot-va-attn.0 likes146 downloads2mo agoHugging Face27AlexMaclean /all-deletion-compressionstext100K<n<1M1 likes142 downloads5y agoHugging Face28AlexMaclean /wikipedia-deletion-compressionstext1K<n<10K2 likes134 downloads5y agoHugging Face29compressionsavant /fw-edu-cl100k0 likes125 downloads6mo agoHugging Face308Planetterraforming /Parameter-Golf-V17-512Cube-XYZ-Priority-Orbital-Compression Parameter Golf V17 — 512Cube Solution Bank This is an English research-control dataset for OpenAI Parameter Golf work. It is not a replacement for FineWeb and must not be used as a substitute training or validation corpus. FineWeb remains the canonical data path for contest scoring. The dataset captures three things: V17 512Cube routing concepts translated into English. Contest and submission guardrails for legal, reproducible BPB reduction. Screenshot-derived scouting observations… See the full description on the dataset page: https://huggingface.co/datasets/8Planetterraforming/Parameter-Golf-V17-512Cube-XYZ-Priority-Orbital-Compression.0 likes116 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.