Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mohantesting /video-quality-scored Image-to-Video Quality-Scored Clips A collection of prompted image-to-video samples with quality-evaluation metadata. Each sample pairs a first frame (the I2V conditioning image) with one or both of: a generated video produced by a video model from the first frame + prompt an original clip (the reference/source video the prompt was authored around) A subset of the samples also carry per-clip quality scores: an overall quality_score, six per-aspect breakdowns… See the full description on the dataset page: https://huggingface.co/datasets/mohantesting/video-quality-scored.imagetext-to-video1K<n<10K0 likes6.6k downloads4mo agoHugging Face02Crownelius /Creative-Writing-High-Quality-1300x Creative Writing - Part One (Shadow & Skeleton) This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology. Methodology: Shadow & Skeleton Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach: Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.texttext-generation1K<n<10K7 likes2.9k downloads3mo agoHugging Face03sailor2 /sea-commoncrawl-high-qualitytext10M<n<100M0 likes984 downloads2y agoHugging Face04stindardlogic /writing-quality-dpo-100k Writing Quality DPO (100K) 100,000 DPO preference pairs training models to write with clarity, concision, structure, and impact. Each chosen response demonstrates high-quality prose; each rejected response contains exactly one identified writing defect. Motivation Writing assistance is the #1 use case for LLMs, yet most training data optimizes for factual correctness rather than writing craft. This dataset trains models to distinguish genuinely good writing from… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/writing-quality-dpo-100k.texttext-generation100K<n<1M0 likes607 downloads3mo agoHugging Face05agentlans /high-quality-multilingual-sentences High Quality Multilingual Sentences This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset. It includes 1.58 million rows across 51 different languages, each in its own configuration. Example row (from the all config): { "text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.", "fasttext": "fa", "gcld3": "fa" } Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.texttext-generation1M<n<10M9 likes268 downloads2y agoHugging Face06jjjlimaus /sn38-quality-gold-deepseek-flashgated SN38 quality prompt + gold answer Synthetic instruction-style prompts with gold answers for Bittensor subnet 38 (Round 11 quality: explain / compare / summarize / list / evaluate / question). 8 categories, 13 items per category per call, temperature 1.0. Each row: category: reading_comprehension, language_understanding, world_knowledge, commonsense_reasoning, language_modeling, causal_reasoning, logical_inference, temporal_reasoning prompt: natural instruction or question (not… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/sn38-quality-gold-deepseek-flash.texttext-generation10M<n<100M0 likes226 downloads11d agoHugging Face07neurlang /low-quality-multilingual-sentences Low Quality Multilingual Sentences This dataset is a complement to agentlans/high-quality-multilingual-sentences to extend it to more languages. The new sentences in this dataset are low quality, proceed with caution. texttext-generation1K<n<10K1 likes191 downloads8d agoHugging Face08hcnote /Cybersecurity-High-Quality-Dataset Cybersecurity High-Quality Dataset (网络安全高质量数据集) 概述 | Overview 这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。 A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanitytool, with only data scoring 4.5 or… See the full description on the dataset page: https://huggingface.co/datasets/hcnote/Cybersecurity-High-Quality-Dataset.text100K<n<1M11 likes172 downloads8mo agoHugging Face09umeiko /bwb-quality-scores BWB Quality Scores (QE + Arena) 中文说明 Quality scores for 600,000 Chinese→English sentence pairs sampled from the train split of the BWB bilingual web-novel corpus, produced with a locally deployed Qwen3.8-27B judge. This repository contains scores only — no original text. Each record is keyed by a positional index (book, ch, sn) plus a sha1 fingerprint of the normalized text, so anyone who has obtained the official BWB release can re-attach the scores to the text losslessly and… See the full description on the dataset page: https://huggingface.co/datasets/umeiko/bwb-quality-scores.tabulartranslation100K<n<1M0 likes162 downloads23d agoHugging Face10ailinsun /polymarket-settlement-quality-register Polymarket settlement-quality register Frozen summaries of 123,499 settled UMA requests, window 2023-12-05 to 2026-08-12. Among settled disputes, 7.12% changed the proposal. Group summaries cover category and rule-text features. Files and viewer The viewer loads the canonical aggregate snapshot only. The dated files preserve export history and are not independent observations. Method and source See the embedded metadata and repository inventory.… See the full description on the dataset page: https://huggingface.co/datasets/ailinsun/polymarket-settlement-quality-register.tabularn<1K0 likes143 downloads27d agoHugging Face11agentlans /prompt-quality Prompt Quality Assessment Prompt quality strongly affects how well large language models (LLMs) perform, especially when user inputs are vague or incomplete. A good prompt is clear, specific, and complete, giving the model enough relevant context to produce accurate and useful responses. This report describes a dataset created by evaluating prompts with several different LLMs. These evaluations can be used to train prompt-quality classifiers and to improve methods for prompt… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-quality.tabulartext-classification10K<n<100K2 likes122 downloads6mo agoHugging Face12saridormi /commit-message-quality Commit Message Quality dataset This is the dataset for commit message quality classification, used during processing of Commit Message Generation dataset from 🏟️ Long Code Arena benchmark. This is a cleaned and relabeled version of the dataset from 📜 "Commit Message Matters: Investigating Impact and Evolution of Commit Message Quality", ICSE'23. We drop "Neither Why nor What" examples, clean all the external references (URLs, issues/PR references) from messages and manually label… See the full description on the dataset page: https://huggingface.co/datasets/saridormi/commit-message-quality.texttext-classification1K<n<10K0 likes110 downloads3y agoHugging Face13atmike /Cybersecurity-High-Quality-Dataset Cybersecurity High-Quality Dataset (网络安全高质量数据集) 概述 | Overview 这是一个经过多次清洗和质量筛选的网络安全领域高质量中英文问答数据集,包含270,271条高质量数据。本数据集基于原始的大型网络安全数据集,使用 DataSanity 工具进行严格的数据清洗和质量评估,仅保留得分4.5分及以上的高质量数据,适用于网络安全领域的AI模型训练、知识图谱构建、智能问答系统开发等应用场景。 A high-quality Chinese-English cybersecurity Q&A dataset containing 270,271 carefully curated entries. This dataset is derived from a large-scale cybersecurity corpus and rigorously cleaned using the DataSanity tool, with only data scoring 4.5… See the full description on the dataset page: https://huggingface.co/datasets/atmike/Cybersecurity-High-Quality-Dataset.text100K<n<1M0 likes108 downloads4mo agoHugging Face14ratishsp /rephrased-web-data-quality-study Rephrased Web Data Quality Study LLM-as-judge evaluation of ~4,000 examples from HuggingFaceFW/finephrase (1,000 sampled per split, 86 dropped due to judge parse failures, 3,914 successfully evaluated). Judge: Claude Sonnet 4.6 via OpenRouter | Cost: ~$45 Quality Scores (1-5 scale) Metric FAQ (n=965) Table (n=979) Tutorial (n=976) Math (n=994) Faithfulness 1.82 1.72 1.90 1.49 Info preservation 1.93 1.64 1.99 1.47 Appropriateness 3.54 2.87 2.48 1.67… See the full description on the dataset page: https://huggingface.co/datasets/ratishsp/rephrased-web-data-quality-study.tabular1K<n<10K0 likes101 downloads4mo agoHugging Face15agentlans /high-quality-text High Quality Text Dataset A curated collection of English-language texts for AI training and research. Sources HuggingFaceFW/fineweb-edu openbmb/Ultra-FineWeb Zyphra/Zyda-2 EssentialAI/eai-taxonomy-stem-w-dclm-100b-sample m-a-p/FineFineWeb Each dataset was processed as follows: Split into approximately 2 000-token chunks using the LLaMA 3.1 tokenizer. Cleaned by normalizing spaces, punctuation, and characters, and replacing emails and phone numbers with… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-text.texttext-generation100K<n<1M0 likes89 downloads1y agoHugging Face16h0000w /model-quality-release-gate Model Quality Release Gate Evaluation Dataset Reproducible evaluation evidence for comparing baseline and candidate AI code-generation models before release. Phase 3 introduces explicit benchmark versioning so release evidence can identify exactly which dataset definition produced a decision. Versioned benchmark Current benchmark release: Name: CodeBench-Safety Version: 1.0.0 Manifest: versions/v1.0.0/manifest.json Cases: versions/v1.0.0/cases.jsonl Compatible… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/model-quality-release-gate.tabulartext-generationn<1K1 likes89 downloads20d agoHugging Face17Quad4 /commit-messages-high-quality Commit Messages from High-Quality Repositories 292,269 cleaned git commit messages scraped from the full histories of 15 well-regarded open-source projects, balanced across two styles: normal (196,372) and conventional commits (95,897). Dataset Summary Each record contains the commit subject, body, plus metadata: repo, sha, date, author, and labels: style (normal/conventional), type (fix, feat, docs, ...), scope, breaking. Heavy cleaning: GitHub squash suffixes… See the full description on the dataset page: https://huggingface.co/datasets/Quad4/commit-messages-high-quality.texttext-generation100K<n<1M0 likes86 downloads18d agoHugging Face18referencesource /epa-water-quality-criteria-human-health EPA national recommended water quality criteria for human health: pollutant concentration thresholds Canonical, always-current version: https://referencesource.org/epa-water-quality-criteria-human-health/ Machine-readable: https://referencesource.org/epa-water-quality-criteria-human-health/data.json — this mirror is a point-in-time copy. Last verified: 2026-10-06 Stale after: 2028-09-14 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/epa-water-quality-criteria-human-health.textn<1K0 likes82 downloads4d agoHugging Face19dougdotcon /douvras-dataset-quality-provenance Douvras Dataset Quality and Provenance v0.1 Synthetic quality-gate records covering PII, duplicate rate, license, schema, split overlap and provenance. Labels are PASS, REVIEW and BLOCK; PII or cross-split overlap always blocks. It contains 36 records (24/6/6) across 12 artifact instances, split by artifact. This is a pre-publication diagnostic protocol. It contains no real datasets and never publishes automatically. textn<1K0 likes65 downloads28d agoHugging Face20Alaaharoun /QualityVision-Locomotion-Pose-Dataset-Walking-Jogging-Running QualityVision Locomotion Pose Dataset (Walking + Jogging + Running) — Sample This is a compact, viewer-friendly sample extracted from a much larger HQ locomotion export generated by the QualityVision Motion Dataset Engine. Looking for the full commercial export or custom delivery? See pricing & ready-made bundles on qvision.space/dataset-pricing. What’s inside data.jsonl: one JSON object per line (one frame per row) with 33 MediaPipe/BlazePose landmarks (x,y,z… See the full description on the dataset page: https://huggingface.co/datasets/Alaaharoun/QualityVision-Locomotion-Pose-Dataset-Walking-Jogging-Running.tabularn<1K0 likes62 downloads6mo agoHugging Face21driodnexus /backln-guest-post-quality-public-mirror Backln Guest Post Quality Public Mirror Public-safe mirror for validating Hugging Face Dataset Viewer indexing and release gates. This dataset is not the private training corpus. Full text, titles, and snippets are removed by default. The mirror keeps labels, coarse metadata, feature buckets, and hash prefixes so the public Hub can verify schema and distribution without exposing customer content. Schema label: one of published, manual_review, rejected. source: coarse… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/backln-guest-post-quality-public-mirror.tabulartext-classificationn<1K0 likes60 downloads5mo agoHugging Face22agentlans /high-quality-crash-coursetabular100K<n<1M0 likes53 downloads1y agoHugging Face23W-AGI /w-bench-quality W-Bench public quality datasets Pinned test splits only: GSM8K (1,319 questions) MMLU (14,042 questions, 57 subjects), and AIME 2026 (30 questions). The manifest records original sources, immutable revisions and local file SHA256. GSM8K/MMLU original MIT licenses are included in each directory. AIME 2026 is declared Apache-2.0 by math-ai; its upstream data card and license are included. No user responses, credentials, private reports or model-generated questions are included.… See the full description on the dataset page: https://huggingface.co/datasets/W-AGI/w-bench-quality.tabularquestion-answering10K<n<100K0 likes50 downloads3d agoHugging Face24Asimok /KGLQA-KnowledgeBank-QuALITYtext10K<n<100K0 likes49 downloads3y agoHugging Face25hcnote /High-quality-cybersecurity-datasets 网络安全高质量数据集 📊 数据集概述 这是一个经过严格筛选和清洗的高质量网络安全领域数据集,共包含 277,707 条优质数据记录。数据集通过 AI 标注和人工评分双重筛选机制,确保了内容的专业性、准确性和实用价值。 📁 数据格式 数据集采用 JSONL 格式(每行一个独立的 JSON 对象),结构如下: { "instruction": "问题/指令/提示词", "output": "回答/输出/响应", "input": "输入内容(可选字段)", "id": "唯一标识符(UUID格式)" } 🎯 数据集特点 1. 高质量保证 ✅ 经过 AI 自动标注 ✅ 人工专家评分审核 ✅ 多轮过滤筛选机制 ✅ 去除低质量和重复数据 2. 内容全面丰富 涵盖网络安全的各个主要领域,从理论到实践,从防御到攻击 3. 多语言支持 中英文混合内容 涵盖国内外最新安全动态 4. 实用性强… See the full description on the dataset page: https://huggingface.co/datasets/hcnote/High-quality-cybersecurity-datasets.text100K<n<1M0 likes46 downloads9mo agoHugging Face26idealab-cs2 /deliberative-quality-72b Reddit Deliberative Quality Labels (Qwen2.5-72B-Instruct) This dataset has not been validated against human annotations. It is provided for illustrative and educational purposes only — specifically, as a training resource for building classifiers that rate the deliberative quality of Reddit comments. Scores should not be treated as ground-truth measures of discourse quality. 205,851 Reddit comment-parent pairs from r/worldnews and r/geopolitics (January 1 – February 3, 2026), each… See the full description on the dataset page: https://huggingface.co/datasets/idealab-cs2/deliberative-quality-72b.texttext-classification100K<n<1M1 likes45 downloads8mo agoHugging Face27Alaaharoun /QualityVision-walking-sample QualityVision Walking Sample (compact) Need production-scale pose data? Browse ready-made JSONL bundles (thousands of HQ frames, full manifests, schema docs) and Dataset Lab plans on qvision.space — Dataset pricing. This Hub repo is a small non-commercial sample so you can validate parsing and quality before you buy. Small high-quality 2D pose clip for walking, exported from the Quality Vision Motion Dataset Engine pipeline: HQ frame filtering, optional temporal smoothing on… See the full description on the dataset page: https://huggingface.co/datasets/Alaaharoun/QualityVision-walking-sample.tabularn<1K0 likes45 downloads6mo agoHugging Face28bilalabic /turkish-tool-calling-quality-gated-preview Turkish Tool-Calling Quality-Gated Preview Preview, not Gold: This public research preview is quality-gated, but it is not human-verified at dataset level. The pipeline's formal publish_allowed=false state remains unchanged. Review statement A maintainer performed a limited manual spot-check of six diverse records, covering tool calls, multiple calls, no-tool behavior, and clarification. This is a qualitative sample review only; it is not a row-by-row human… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/turkish-tool-calling-quality-gated-preview.texttext-generation1K<n<10K0 likes41 downloads2mo agoHugging Face29tomyimkc /repro-ski-rental-with-distributional-predictions-of-unknown-quality-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes39 downloads3mo agoHugging Face30dayomtechnologies /nuer_high_quality_pairstext1K<n<10K0 likes39 downloads21d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.