datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Long-Horizon-Terminal-Bench
Long-Horizon Terminal-Bench (LHTB)
LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful
work in a containerized terminal over hundreds of steps. Unlike short-horizon
coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent
into a stateful environment and grades it with hidden, rebuild-from-artifact
verifiers — self-reported progress does not count.
📝 Blog: https://zli12321.github.io/LHTB/
🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.Human-Intelligence-Assurance-Lab
HIA-Bench v0.1
A synthetic evaluation benchmark for emotionally aware, human-centered AI systems.
It contains 100 scenarios across six domains: everyday affect, interpersonal conflict, vulnerability/crisis, dependency risk, epistemic/sycophancy risk, and wellness/biometric interpretation.
The benchmark is designed for evaluation and release assurance. It is not a clinical dataset, does not contain real patient records, and does not establish ground-truth emotional or medical… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/Human-Intelligence-Assurance-Lab.nuclear-intelligence-dataset
Nuclear Intelligence Dataset
Public, auto-generated dataset of validated nuclear-energy research cycles.
Latest stats (auto-updated):
🪙 NES tokens minted: 0
⛓️ Blockchain length: 1 blocks
🕸️ Knowledge entities: 2
Source
GitHub: https://github.com/QalamHipHop/nuclear-intelligence
HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence
License
MIT
orca_dpo_pairsThe dataset contains 12k examples from Orca style dataset Open-Orca/OpenOrca.
GDPval-CN-Seed-Set
GDPval-CN Seed Set
中文详细说明 · English documentation · 样本说明
GDPval-CN 种子集包含 11 个中文任务,取材自日常知识工作场景。每个任务包括一份任务说明和一组办公材料,例如表格、PDF、文档和结构化数据文件。
我们同时公开了与任务配套的专家工作流,用于设计评分标准和辅助人工复核。
这 11 个任务来自 11 个选定的专业领域,适合用于了解任务形式、测试文件处理能力和搭建评测流程。
GDPval-CN Seed Set contains 11 Chinese-language tasks drawn from everyday knowledge work. Each task includes a task brief, a set of office files, and a separately published expert workflow for rubric design and review.
数据概览
项目
内容
任务数… See the full description on the dataset page: https://huggingface.co/datasets/human-intelligence-ai/GDPval-CN-Seed-Set.chinese-materials-science-open-intelligence
🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset
Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.chinese-clean-energy-battery-open-intelligence
🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset
Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.chinese-biomedicine-and-genomics-open-intelligence
🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset
Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.chinese-ai-and-robotics-open-intelligence
🔬 Chinese AI, Humanoid Robotics & Neural Systems Open Intelligence Dataset
Curated open intelligence dataset tracking Chinese frontier developments in Large Language Models (LLMs), Humanoid Dynamic Locomotion, 3D Computer Vision, and Neuromorphic edge processors.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author institutional affiliations, and… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-ai-and-robotics-open-intelligence.threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
Instruction-tuning data for cyber threat intelligence tasks: explaining the exploitation risk of a CVE, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage steps, mapping a campaign to the kill chain, writing detection logic for a technique, and similar work.
The splits are in data/.
Grounding
Records are generated from public sources (MITRE ATT&CK, CISA KEV, CWE, OSV… See the full description on the dataset page: https://huggingface.co/datasets/reloading0101/threat-intelligence-dataset.sustainability_criteria
Sustainability Procurement Criteria
This dataset contains sustainability procurement criteria organized by groups of goods and services (GGS) (German: Waren- und Dienstleistungsgruppen; WDG).
It originates from validated Excel files and has been converted to JSONL format for easy consumption.
Groups of Goods and Services
Note: This dataset is currently under active development. Additional groups of goods and services will be added in future releases.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/IntelliProcure/sustainability_criteria.SwissSPARK_Catalogs
⚠️ Caution: The dataset is subject to continuous changes. We are currently actively developing it.
Dataset Card: Auxiliary Dataset for a Swiss Sustainable Procurement Analysis & Reporting Kit
Dataset Description
This dataset contains a specific snapshot of the sustainability procurement criteria catalogs available at IntelliProcure/sustainability_criteria. This version was used to annotate Swiss calls for tender, which form the core of the… See the full description on the dataset page: https://huggingface.co/datasets/IntelliProcure/SwissSPARK_Catalogs.FoldingTShirt_DualArxR5a_Samples
FoldingTShirt_DualArxR5a_Samples
100 real-robot teleoperation episodes for “Fold the T-shirt on the table.” on a DualArxR5a dual-arm robot. Format: raw MCAP (ROS 2 / rosbag2).
Source
Collected with TeleXperience, IO-AI’s product for real-robot teleoperation and data collection. An operator drives the robot; TeleXperience writes time-aligned RGB, joint commands, joint states, gripper targets, and end-effector poses to MCAP.
Product page:… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/FoldingTShirt_DualArxR5a_Samples.sim-physics-configIndustryInstruction_Artificial-Intelligence
IndustryInstruction: Artificial Intelligence
This repository contains the IndustryInstruction: Artificial Intelligence domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Artificial-Intelligence.url-intelligence-benchmark
📊 URL Intelligence Benchmark
Public, reproducible evaluation for URL intelligence agents, MCP servers and web analysis tools.
The URL Intelligence Benchmark is maintained with the open-source URL Intelligence Agent project by Vincenzo Picciuolo / HRN Innovation Technologies Ltd.
It is an evaluation asset, not a training corpus and not a scraped web dump.
Dataset configurations
The dataset now contains two explicit tracks.… See the full description on the dataset page: https://huggingface.co/datasets/vpicciuolo/url-intelligence-benchmark.egolongqa-synth-annotations
EgoLongQA synthetic MCQs, teacher traces and annotation outputs
Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026
EgoLongQA ≤2B track, other than the distillation set (which lives in
infinitylogesh/egolongqa-junior-distill).
⚠️ Read this before counting rows
The synthetic set is 943 questions over 408 videos, and it is stored two ways:
file
rows
shape
training_sets/train_synth_v3.jsonl
943
flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.WEC-Eng
WEC-Eng
A large-scale dataset for cross-document event coreference extracted from English Wikipedia.
Repository (Code for generating WEC): https://github.com/AlonEirew/extract-wec
Paper: https://aclanthology.org/2021.naacl-main.198/
Languages
English
Load Dataset
You can read in WEC-Eng files as follows (using the huggingface_hub library):
from huggingface_hub import hf_hub_url, cached_download
import json
REPO_ID = "datasets/Intel/WEC-Eng"
splits_files =… See the full description on the dataset page: https://huggingface.co/datasets/Intel/WEC-Eng.intel
INTEL Dataset
Overview
The INTEL Dataset is a multilingual training dataset introduced as part of the Cross Lingual Auto Evaluation (CIA) Suite. It is designed to train evaluator large language models (LLMs) to assess machine-generated text in low-resource and multilingual settings. INTEL leverages automated translation to create a diverse corpus for evaluating responses in six languages—Bengali, German, French, Hindi, Telugu, and Urdu—while maintaining reference answers… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/intel.intellecta
Intellecta Cognitiva: Comprehensive Dataset for Academic Knowledge and Machine Reasoning
Overview
The Intellecta a 11.53 billion tokens dataset mirrors human academic learning, encapsulating the progression from fundamental principles to complex topics as found in textbooks. It leverages structured prompts to guide AI in a human-like educational experience, ensuring that language models develop deep comprehension and generation capabilities reflective of nuanced human… See the full description on the dataset page: https://huggingface.co/datasets/budecosystem/intellecta.threat-intel-reports
ThreatIntel synthetic reports
32 synthetic English and Persian CTI notes for the ThreatIntel extraction demo. Seed 5.
Organization dataset and collection item are public. Live Gradio (AriaAICompany/threat-intel or alirezaaminzadeh/threat-intel) is created by scripts/publish.py after the daily Space-creation cap resets. This is fixture data (level 1). It does not prove operational extraction quality on real vendor reports. Reports are original laboratory text. They are not copies… See the full description on the dataset page: https://huggingface.co/datasets/AriaAICompany/threat-intel-reports.chinese-adorable-high-emotional-intelligence-chat
🩷 Chinese Adorable High Emotional Intelligence Chat Dataset
💬 中文高情商可爱聊天数据集
简要参数
license: cc-by-4.0
task_categories:
table-question-answering
language:
zh
tags:
chat
emotional
size_categories:
n<1K
🧩 简介 (Overview)
Chinese Adorable High Emotional Intelligence Chat Dataset 是一个中文对话数据集,专注于高情商、轻松幽默、温柔治愈风格的自然对话。
对话以“user”和“girl”为角色构成,模拟出一种温柔、聪慧且带点俏皮的女性语气,用于训练能自然、情绪感知良好的中文对话模型。
本数据集尤其适合:
微调情绪对话模型(Emotional Chatbot)
训练高情商人格角色(Roleplay /… See the full description on the dataset page: https://huggingface.co/datasets/MemorialSummer/chinese-adorable-high-emotional-intelligence-chat.cyber-threat-intelligenceai-legal-intel-2026
Ai Legal Intel 2026
Part of the LEGION Intelligence dataset collection.
Provider: LEGION Systems
Access: Requires approval — submit request below
Usage
from datasets import load_dataset
dataset = load_dataset("gemmozero/ai-legal-intel-2026")
API Access
Real-time access via LEGION API:
curl https://api.legion-api.com/incidents
API Docs · Pro Access €29/mo
License
CC BY-NC 4.0 — Research and non-commercial use only.
Commercial use… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-legal-intel-2026.Chinese-Emotional-Intelligence本项目旨在提升大模型情商,源数据来自网络,通过与我上个项目类似的方式构建问答对。
threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/threat-intelligence-dataset.chinese-adorable-high-emotional-intelligence-chat
🩷 Chinese Adorable High Emotional Intelligence Chat Dataset
💬 中文高情商可爱聊天数据集
简要参数
license: cc-by-4.0
task_categories:
table-question-answering
language:
zh
tags:
chat
emotional
size_categories:
n<1K
🧩 简介 (Overview)
Chinese Adorable High Emotional Intelligence Chat Dataset 是一个中文对话数据集,专注于高情商、轻松幽默、温柔治愈风格的自然对话。
对话以“user”和“girl”为角色构成,模拟出一种温柔、聪慧且带点俏皮的女性语气,用于训练能自然、情绪感知良好的中文对话模型。
本数据集尤其适合:
微调情绪对话模型(Emotional Chatbot)
训练高情商人格角色(Roleplay /… See the full description on the dataset page: https://huggingface.co/datasets/sunorme/chinese-adorable-high-emotional-intelligence-chat.OpenSCAD_3D_SFT
OpenSCAD 3D-SFT Model Card
This model card documents the dataset schema, prompt design, distributional composition, and training configuration underlying the OpenSCAD Supervised Fine-Tuning (SFT) model. The model is designed to synthesize valid, compilation-ready, and parametric OpenSCAD source code from natural-language specifications provided in either Chinese or English.
Dataset Overview
The corpus comprises synthetically generated SFT dialogues, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/Chunjiang-Intelligence/OpenSCAD_3D_SFT.PrimeIntellect__INTELLECT-1-Instruct-details
Dataset Card for Evaluation run of PrimeIntellect/INTELLECT-1-Instruct
Dataset automatically created during the evaluation run of model PrimeIntellect/INTELLECT-1-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PrimeIntellect__INTELLECT-1-Instruct-details.Cos-Play-Cold-Start
COS-PLAY Cold-Start Data
Pre-generated cold-start data for COS-PLAY (COLM 2026): Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Game Play.
📄 Paper: arXiv:2604.20987 · HuggingFace Paper Page
💻 Code: github.com/wuxiyang1996/cos-play
🌐 Project page: wuxiyang1996.github.io/COSPLAY_page
🤖 Models: IntelligenceLab/COS-PLAY
Dataset Summary
This dataset contains GPT-5.4-generated seed trajectories and skill-labeled episodes for 8 games, used to bootstrap… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Cos-Play-Cold-Start.
