datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
video-quality-scored
Image-to-Video Quality-Scored Clips
A collection of prompted image-to-video samples with quality-evaluation metadata.
Each sample pairs a first frame (the I2V conditioning image) with one or both
of:
a generated video produced by a video model from the first frame + prompt
an original clip (the reference/source video the prompt was authored around)
A subset of the samples also carry per-clip quality scores: an overall
quality_score, six per-aspect breakdowns… See the full description on the dataset page: https://huggingface.co/datasets/mohantesting/video-quality-scored.Creative-Writing-High-Quality-1300x
Creative Writing - Part One (Shadow & Skeleton)
This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology.
Methodology: Shadow & Skeleton
Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach:
Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.sea-commoncrawl-high-qualitywriting-quality-dpo-100k
Writing Quality DPO (100K)
100,000 DPO preference pairs training models to write with clarity, concision, structure, and impact. Each chosen response demonstrates high-quality prose; each rejected response contains exactly one identified writing defect.
Motivation
Writing assistance is the #1 use case for LLMs, yet most training data optimizes for factual correctness rather than writing craft. This dataset trains models to distinguish genuinely good writing from… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/writing-quality-dpo-100k.high-quality-multilingual-sentences
High Quality Multilingual Sentences
This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset.
It includes 1.58 million rows across 51 different languages, each in its own configuration.
Example row (from the all config):
{
"text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.",
"fasttext": "fa",
"gcld3": "fa"
}
Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.sn38-quality-gold-deepseek-flash
SN38 quality prompt + gold answer
Synthetic instruction-style prompts with gold answers for Bittensor subnet 38
(Round 11 quality: explain / compare / summarize / list / evaluate / question).
8 categories, 13 items per category per call, temperature 1.0.
Each row:
category: reading_comprehension, language_understanding, world_knowledge, commonsense_reasoning, language_modeling, causal_reasoning, logical_inference, temporal_reasoning
prompt: natural instruction or question (not… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/sn38-quality-gold-deepseek-flash.low-quality-multilingual-sentences
Low Quality Multilingual Sentences
This dataset is a complement to agentlans/high-quality-multilingual-sentences to extend it to more languages.
The new sentences in this dataset are low quality, proceed with caution.
Cybersecurity-High-Quality-Dataset
Cybersecurity High-Quality Dataset (网络安全高质量数据集)
概述 | Overview
这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。
A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanitytool, with only data scoring 4.5 or… See the full description on the dataset page: https://huggingface.co/datasets/hcnote/Cybersecurity-High-Quality-Dataset.bwb-quality-scores
BWB Quality Scores (QE + Arena)
中文说明
Quality scores for 600,000 Chinese→English sentence pairs sampled from the train split of the BWB bilingual web-novel corpus, produced with a locally deployed Qwen3.8-27B judge.
This repository contains scores only — no original text. Each record is keyed by a positional index (book, ch, sn) plus a sha1 fingerprint of the normalized text, so anyone who has obtained the official BWB release can re-attach the scores to the text losslessly and… See the full description on the dataset page: https://huggingface.co/datasets/umeiko/bwb-quality-scores.polymarket-settlement-quality-register
Polymarket settlement-quality register
Frozen summaries of 123,499 settled UMA requests, window 2023-12-05 to 2026-08-12. Among settled disputes, 7.12% changed the proposal. Group summaries cover category and rule-text features.
Files and viewer
The viewer loads the canonical aggregate snapshot only. The dated files preserve export history and are not independent observations.
Method and source
See the embedded metadata and repository inventory.… See the full description on the dataset page: https://huggingface.co/datasets/ailinsun/polymarket-settlement-quality-register.prompt-quality
Prompt Quality Assessment
Prompt quality strongly affects how well large language models (LLMs) perform, especially when user inputs are vague or incomplete. A good prompt is clear, specific, and complete, giving the model enough relevant context to produce accurate and useful responses.
This report describes a dataset created by evaluating prompts with several different LLMs. These evaluations can be used to train prompt-quality classifiers and to improve methods for prompt… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-quality.commit-message-quality
Commit Message Quality dataset
This is the dataset for commit message quality classification, used during processing of Commit Message Generation dataset from
🏟️ Long Code Arena benchmark.
This is a cleaned and relabeled version of the dataset from 📜 "Commit Message Matters: Investigating Impact and Evolution of Commit Message Quality", ICSE'23. We drop "Neither Why nor What" examples, clean all the external references (URLs, issues/PR references) from messages and manually label… See the full description on the dataset page: https://huggingface.co/datasets/saridormi/commit-message-quality.Cybersecurity-High-Quality-Dataset
Cybersecurity High-Quality Dataset (网络安全高质量数据集)
概述 | Overview
这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。
A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanity tool, with only data scoring 4.5… See the full description on the dataset page: https://huggingface.co/datasets/atmike/Cybersecurity-High-Quality-Dataset.rephrased-web-data-quality-study
Rephrased Web Data Quality Study
LLM-as-judge evaluation of ~4,000 examples from HuggingFaceFW/finephrase (1,000 sampled per split, 86 dropped due to judge parse failures, 3,914 successfully evaluated).
Judge: Claude Sonnet 4.6 via OpenRouter | Cost: ~$45
Quality Scores (1-5 scale)
Metric
FAQ (n=965)
Table (n=979)
Tutorial (n=976)
Math (n=994)
Faithfulness
1.82
1.72
1.90
1.49
Info preservation
1.93
1.64
1.99
1.47
Appropriateness
3.54
2.87
2.48
1.67… See the full description on the dataset page: https://huggingface.co/datasets/ratishsp/rephrased-web-data-quality-study.high-quality-text
High Quality Text Dataset
A curated collection of English-language texts for AI training and research.
Sources
HuggingFaceFW/fineweb-edu
openbmb/Ultra-FineWeb
Zyphra/Zyda-2
EssentialAI/eai-taxonomy-stem-w-dclm-100b-sample
m-a-p/FineFineWeb
Each dataset was processed as follows:
Split into approximately 2 000-token chunks using the LLaMA 3.1 tokenizer.
Cleaned by normalizing spaces, punctuation, and characters, and replacing emails and phone numbers with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-text.model-quality-release-gate
Model Quality Release Gate Evaluation Dataset
Reproducible evaluation evidence for comparing baseline and candidate AI code-generation models before release.
Phase 3 introduces explicit benchmark versioning so release evidence can identify exactly which dataset definition produced a decision.
Versioned benchmark
Current benchmark release:
Name: CodeBench-Safety
Version: 1.0.0
Manifest: versions/v1.0.0/manifest.json
Cases: versions/v1.0.0/cases.jsonl
Compatible… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/model-quality-release-gate.commit-messages-high-quality
Commit Messages from High-Quality Repositories
292,269 cleaned git commit messages scraped from the full histories of 15 well-regarded open-source
projects, balanced across two styles: normal (196,372) and
conventional commits (95,897).
Dataset Summary
Each record contains the commit subject, body, plus metadata: repo, sha, date,
author, and labels: style (normal/conventional), type (fix, feat, docs, ...),
scope, breaking.
Heavy cleaning: GitHub squash suffixes… See the full description on the dataset page: https://huggingface.co/datasets/Quad4/commit-messages-high-quality.epa-water-quality-criteria-human-health
EPA national recommended water quality criteria for human health: pollutant concentration thresholds
Canonical, always-current version: https://referencesource.org/epa-water-quality-criteria-human-health/
Machine-readable: https://referencesource.org/epa-water-quality-criteria-human-health/data.json — this mirror is a point-in-time copy.
Last verified: 2026-10-06
Stale after: 2028-09-14 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/epa-water-quality-criteria-human-health.douvras-dataset-quality-provenance
Douvras Dataset Quality and Provenance v0.1
Synthetic quality-gate records covering PII, duplicate rate, license, schema,
split overlap and provenance. Labels are PASS, REVIEW and BLOCK; PII or
cross-split overlap always blocks. It contains 36 records (24/6/6) across 12
artifact instances, split by artifact.
This is a pre-publication diagnostic protocol. It contains no real datasets and
never publishes automatically.
QualityVision-Locomotion-Pose-Dataset-Walking-Jogging-Running
QualityVision Locomotion Pose Dataset (Walking + Jogging + Running) — Sample
This is a compact, viewer-friendly sample extracted from a much larger HQ locomotion export generated by the QualityVision Motion Dataset Engine.
Looking for the full commercial export or custom delivery? See pricing & ready-made bundles on qvision.space/dataset-pricing.
What’s inside
data.jsonl: one JSON object per line (one frame per row) with 33 MediaPipe/BlazePose landmarks (x,y,z… See the full description on the dataset page: https://huggingface.co/datasets/Alaaharoun/QualityVision-Locomotion-Pose-Dataset-Walking-Jogging-Running.backln-guest-post-quality-public-mirror
Backln Guest Post Quality Public Mirror
Public-safe mirror for validating Hugging Face Dataset Viewer indexing and release gates. This dataset is not the private training corpus.
Full text, titles, and snippets are removed by default. The mirror keeps labels, coarse metadata, feature buckets, and hash prefixes so the public Hub can verify schema and distribution without exposing customer content.
Schema
label: one of published, manual_review, rejected.
source: coarse… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/backln-guest-post-quality-public-mirror.high-quality-crash-coursew-bench-quality
W-Bench public quality datasets
Pinned test splits only: GSM8K (1,319 questions) MMLU (14,042 questions, 57 subjects), and AIME 2026 (30 questions).
The manifest records original sources, immutable revisions and local file SHA256.
GSM8K/MMLU original MIT licenses are included in each directory. AIME 2026
is declared Apache-2.0 by math-ai; its upstream data card and license are included. No user responses,
credentials, private reports or model-generated questions are included.… See the full description on the dataset page: https://huggingface.co/datasets/W-AGI/w-bench-quality.KGLQA-KnowledgeBank-QuALITYHigh-quality-cybersecurity-datasets
网络安全高质量数据集
📊 数据集概述
这是一个经过严格筛选和清洗的高质量网络安全领域数据集,共包含 277,707 条优质数据记录。数据集通过 AI 标注和人工评分双重筛选机制,确保了内容的专业性、准确性和实用价值。
📁 数据格式
数据集采用 JSONL 格式(每行一个独立的 JSON 对象),结构如下:
{
"instruction": "问题/指令/提示词",
"output": "回答/输出/响应",
"input": "输入内容(可选字段)",
"id": "唯一标识符(UUID格式)"
}
🎯 数据集特点
1. 高质量保证
✅ 经过 AI 自动标注
✅ 人工专家评分审核
✅ 多轮过滤筛选机制
✅ 去除低质量和重复数据
2. 内容全面丰富
涵盖网络安全的各个主要领域,从理论到实践,从防御到攻击
3. 多语言支持
中英文混合内容
涵盖国内外最新安全动态
4. 实用性强… See the full description on the dataset page: https://huggingface.co/datasets/hcnote/High-quality-cybersecurity-datasets.deliberative-quality-72b
Reddit Deliberative Quality Labels (Qwen2.5-72B-Instruct)
This dataset has not been validated against human annotations. It is provided for illustrative and educational purposes only — specifically, as a training resource for building classifiers that rate the deliberative quality of Reddit comments. Scores should not be treated as ground-truth measures of discourse quality.
205,851 Reddit comment-parent pairs from r/worldnews and r/geopolitics (January 1 – February 3, 2026), each… See the full description on the dataset page: https://huggingface.co/datasets/idealab-cs2/deliberative-quality-72b.QualityVision-walking-sample
QualityVision Walking Sample (compact)
Need production-scale pose data? Browse ready-made JSONL bundles (thousands of HQ frames, full manifests, schema docs) and Dataset Lab plans on qvision.space — Dataset pricing. This Hub repo is a small non-commercial sample so you can validate parsing and quality before you buy.
Small high-quality 2D pose clip for walking, exported from the Quality Vision Motion Dataset Engine pipeline: HQ frame filtering, optional temporal smoothing on… See the full description on the dataset page: https://huggingface.co/datasets/Alaaharoun/QualityVision-walking-sample.turkish-tool-calling-quality-gated-preview
Turkish Tool-Calling Quality-Gated Preview
Preview, not Gold: This public research preview is quality-gated, but it
is not human-verified at dataset level. The pipeline's formal
publish_allowed=false state remains unchanged.
Review statement
A maintainer performed a limited manual spot-check of six diverse records,
covering tool calls, multiple calls, no-tool behavior, and clarification. This
is a qualitative sample review only; it is not a row-by-row human… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/turkish-tool-calling-quality-gated-preview.repro-ski-rental-with-distributional-predictions-of-unknown-quality-traces
Agent traces
Agent sessions published from a Trackio Logbook.
nuer_high_quality_pairs
