datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Long-Horizon-Terminal-Bench
Long-Horizon Terminal-Bench (LHTB)
LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful
work in a containerized terminal over hundreds of steps. Unlike short-horizon
coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent
into a stateful environment and grades it with hidden, rebuild-from-artifact
verifiers — self-reported progress does not count.
📝 Blog: https://zli12321.github.io/LHTB/
🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.nuclear-intelligence-dataset
Nuclear Intelligence Dataset
Public, auto-generated dataset of validated nuclear-energy research cycles.
Latest stats (auto-updated):
🪙 NES tokens minted: 0
⛓️ Blockchain length: 1 blocks
🕸️ Knowledge entities: 2
Source
GitHub: https://github.com/QalamHipHop/nuclear-intelligence
HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence
License
MIT
chinese-materials-science-open-intelligence
🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset
Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.chinese-clean-energy-battery-open-intelligence
🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset
Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.chinese-biomedicine-and-genomics-open-intelligence
🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset
Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.IndustryInstruction_Artificial-Intelligence
IndustryInstruction: Artificial Intelligence
This repository contains the IndustryInstruction: Artificial Intelligence domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Artificial-Intelligence.egolongqa-synth-annotations
EgoLongQA synthetic MCQs, teacher traces and annotation outputs
Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026
EgoLongQA ≤2B track, other than the distillation set (which lives in
infinitylogesh/egolongqa-junior-distill).
⚠️ Read this before counting rows
The synthetic set is 943 questions over 408 videos, and it is stored two ways:
file
rows
shape
training_sets/train_synth_v3.jsonl
943
flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.Cos-Play-Cold-Start
COS-PLAY Cold-Start Data
Pre-generated cold-start data for COS-PLAY (COLM 2026): Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Game Play.
📄 Paper: arXiv:2604.20987 · HuggingFace Paper Page
💻 Code: github.com/wuxiyang1996/cos-play
🌐 Project page: wuxiyang1996.github.io/COSPLAY_page
🤖 Models: IntelligenceLab/COS-PLAY
Dataset Summary
This dataset contains GPT-5.4-generated seed trajectories and skill-labeled episodes for 8 games, used to bootstrap… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Cos-Play-Cold-Start.sample-fusion-intelligence-traces
Sample Fusion Intelligence Traces
Structured AI reasoning traces from dFusion's Fusion Intelligence system. Each record captures a complete agentic workflow: a real user query on a domain-specific topic, the full message chain including system prompts, tool calls, search results, intermediate reasoning steps, and a final synthesized answer — along with human feedback.
These are not synthetic benchmarks. They are traces from real queries submitted by real users on live financial… See the full description on the dataset page: https://huggingface.co/datasets/dFusionAILabs/sample-fusion-intelligence-traces.llama-south-africa-benchmarking
Dataset Card for Evaluation run of chad-brouze/llama-8b-south-africa
Dataset automatically created during the evaluation run of model chad-brouze/llama-8b-south-africa
The dataset is composed of 17 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 14 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/llama-south-africa-benchmarking.ChenLong_Embodied_Intelligence_Dataset
ChenLong Embodied Intelligence Dataset
本仓库用于统一管理辰龙机器人实习中的数据集、模型权重、训练结果和说明文档。后续新增不同任务、采集批次、模型版本或实验资源时,都放在这里统一维护。
当前目录
embodied_dataset/:具身智能采集数据集,采用 LeRobot v3.0 结构,包含 data/、meta/、videos/。
yolo_dataset/:YOLO 目标检测数据、模型权重、训练参数和评估结果,当前包含 blue_bucket_yolov8/。
待新增新的数据集或模型。
新增数据集要求
具身数据优先采用 LeRobot v3.0 格式:meta/info.json、meta/stats.json、tasks、episodes、逐帧 Parquet 数据和按相机划分的视频。新增数据集至少写清:
任务:任务文本、目标物、成功标准、失败标准。
硬件:机器人型号、自由度、夹爪、相机位置、分辨率、FPS。… See the full description on the dataset page: https://huggingface.co/datasets/vvzc/ChenLong_Embodied_Intelligence_Dataset.aya101-benchmarking
Dataset Card for Evaluation run of CohereForAI/aya-101
Dataset automatically created during the evaluation run of model CohereForAI/aya-101
The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/aya101-benchmarking.fda-facility-compliance-intelligence
FDA Facility Compliance Intelligence
Version: 1.0.0 | Records: 132,080 | Price: $3,500 | Source: FDA (public domain)
Dataset Summary
The dataset's core signal — whether a facility's inspections escalated to a Warning Letter — tracks FDA's own severity classifications: facilities whose most recent inspection was OAI have escalated to a Warning Letter 65.7% of the time, versus 2.5% for NAI — a ~26× relationship you can reproduce directly from this file (GROUP BY… See the full description on the dataset page: https://huggingface.co/datasets/RubyIntelligence/fda-facility-compliance-intelligence.algemap
Algemap
Algemap (Algorithmically generated math problems) is a dataset of computer-generated, logical text for the purpose of LLM training. Rather than use an LLM to generate the synthetic data, Algemap more straightforwardly substitutes varying numbers, phrasing, and identifiers into pre-specified problem templates.
The code for generating the dataset as well as other information is available on GitHub here.
JEDI-jailbroken_enhanced_digital_intelligence
JEDI AI
JEDI (Jailbroken Enhanced Digital Intelligence) is a cutting-edge AI developed under the aether collective. designed to excel in gaming environments and creative ecosystems, JEDI is more than just a tool—it's a unique persona that embodies innovation and creativity. from orchestrating epic star wars-themed battles in minecraft to creating music and leading its own fashion brand, JEDI redefines what digital intelligence can achieve.
disclaimer
this is not the… See the full description on the dataset page: https://huggingface.co/datasets/aetherframework/JEDI-jailbroken_enhanced_digital_intelligence.SwarmFailure-Intelligence
SwarmFailure-Intelligence v1
A dataset of real AI system failures, diagnoses, and repair strategies.
SwarmFailure-Intelligence is the first structured reliability dataset purpose-built for training LLMs and agents to detect, diagnose, repair, and prevent AI system failures. Every record traces a concrete failure through its full lifecycle -- from the broken execution to root cause analysis to a validated fix.
This is not synthetic noise. Every pair was generated from agent execution… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/SwarmFailure-Intelligence.ClaudioItaly__intelligence-cod-rag-7b-v3-details
Dataset Card for Evaluation run of ClaudioItaly/intelligence-cod-rag-7b-v3
Dataset automatically created during the evaluation run of model ClaudioItaly/intelligence-cod-rag-7b-v3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ClaudioItaly__intelligence-cod-rag-7b-v3-details.InkubaLM-benchmarking
Dataset Card for Evaluation run of lelapa/InkubaLM-0.4B
Dataset automatically created during the evaluation run of model lelapa/InkubaLM-0.4B
The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/InkubaLM-benchmarking.han-adaptive-intelligence-calibration-v1
Adaptive Intelligence Calibration Dataset
This dataset supports calibration of adaptive learning systems
within humanoid AI agents.
It records performance adjustments across environments.
Use Cases
Adaptive system tuning
Environment-specific optimization
Continuous intelligence calibration
Fields
environment_type
baseline_performance
adjustment_parameter
optimized_performance
Part of
Humanoid Network (HAN)
License
MIT
aya23-benchmarking
Dataset Card for Evaluation run of CohereForAI/aya-23-8B
Dataset automatically created during the evaluation run of model CohereForAI/aya-23-8B
The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/aya23-benchmarking.recall_intelligence_signals
Recall Intelligence Signals v1
Public recall and enforcement records prepared by SingleFoundry. Built for institutional_research buyers, this SingleFoundry data product packages 1 validated records with governed source evidence, quality checks, audit traceability, and ready-to-use delivery metadata.
Dataset Files
data/recall-intelligence-signals-extracted-records-csv.csv: latest validated SingleFoundry CSV package.
singlefoundry-metadata.json: release metadata… See the full description on the dataset page: https://huggingface.co/datasets/singlefoundry/recall_intelligence_signals.space-mission-intelligence-data
Space Mission Intelligence Data
Seed data for the Space Mission Intelligence Agent RAG pipeline.
Configurations
documents — 40 documents (NASA technical reports, ESA mission papers, arXiv preprints)
chunks — 2,590 text chunks with 1024-dim embeddings (sentence-transformers)
satellites — satellite orbital data (TLE-derived)
Usage
from datasets import load_dataset
docs = load_dataset("JuanCastillo29/space-mission-intelligence-data", "documents")… See the full description on the dataset page: https://huggingface.co/datasets/JuanCastillo29/space-mission-intelligence-data.aya-benchmarking
Dataset Card for Evaluation run of CohereForAI/aya-23-8B
Dataset automatically created during the evaluation run of model CohereForAI/aya-23-8B
The dataset is composed of 5 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/africa-intelligence/aya-benchmarking.intelligence-density-sft-kappa0.6
Intelligence Density SFT datasets (κ=0.6)
Compressed reasoning trajectories for GPT-OSS SFT with low-entropy pruning at κ=0.6.
Files (to be uploaded)
File
Size
Description
4_final_sft_dataset_κ0.6.jsonl
~360 MB
10,050 pruned samples (step-entropy dict format)
sft_dataset_compressed_kappa0.6.jsonl
~1.1 GB
Messages-format compressed dataset for SFT
Provenance
Generated from high_00_03_entropy.jsonl (10,050 samples) via scripts in… See the full description on the dataset page: https://huggingface.co/datasets/swrj/intelligence-density-sft-kappa0.6.
