datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Long-Horizon-Terminal-Bench
Long-Horizon Terminal-Bench (LHTB)
LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful
work in a containerized terminal over hundreds of steps. Unlike short-horizon
coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent
into a stateful environment and grades it with hidden, rebuild-from-artifact
verifiers — self-reported progress does not count.
📝 Blog: https://zli12321.github.io/LHTB/
🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.nuclear-intelligence-dataset
Nuclear Intelligence Dataset
Public, auto-generated dataset of validated nuclear-energy research cycles.
Latest stats (auto-updated):
🪙 NES tokens minted: 0
⛓️ Blockchain length: 1 blocks
🕸️ Knowledge entities: 2
Source
GitHub: https://github.com/QalamHipHop/nuclear-intelligence
HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence
License
MIT
XL-SafetyBench
XL-SafetyBench
A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
⚠️ Content Warning: This dataset contains adversarial prompts and
culturally sensitive content for safety and cultural-evaluation research.
By using this dataset, you agree to use it solely for research purposes
and not for malicious applications.
Paper: https://arxiv.org/abs/2605.05662
Eval Code: github.com/AIM-Intelligence/XL-SafetyBench
Overview… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/XL-SafetyBench.threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
Instruction-tuning data for cyber threat intelligence tasks: explaining the exploitation risk of a CVE, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage steps, mapping a campaign to the kill chain, writing detection logic for a technique, and similar work.
The splits are in data/.
Grounding
Records are generated from public sources (MITRE ATT&CK, CISA KEV, CWE, OSV… See the full description on the dataset page: https://huggingface.co/datasets/reloading0101/threat-intelligence-dataset.COMPASS-Policy-Alignment-Testbed-Dataset
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs
This dataset evaluates how well Large Language Models (LLMs) follow organization-specific policies in realistic enterprise-style settings.
What is COMPASS?
COMPASS is a framework for evaluating policy alignment: given only an organization’s policy (e.g., allow/deny rules), it enables you to benchmark whether an LLM’s responses comply with that policy in structured, enterprise-like… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/COMPASS-Policy-Alignment-Testbed-Dataset.Arc-ATLAS-Teach-v1
Arc-ATLAS-Teach
Summary
This revision bundles 624 high-quality adaptive teaching examples that were generated and validated with the latest five-pass pipeline. Every dialogue walks through the full instructional arc—probe, draft plan, checkpoint feedback, revised plan, and final solution—so the teaching policy observes the complete adjustment process without ever seeing the canonical answer. Probe turns capture the student’s diagnostic attempt, teacher plans and… See the full description on the dataset page: https://huggingface.co/datasets/Arc-Intelligence/Arc-ATLAS-Teach-v1.unified-vulnerability-intelligence-dataset
Unified Vulnerability Intelligence Dataset (UVID) v3.0 — Cyber Security Knowledge Graph
UVID is a structured cyber security knowledge graph that unifies multiple
vulnerability classification frameworks into a single knowledge base. Each of the
250 records describes one application/software security vulnerability and links
it — where authoritative data exists — across CWE, CAPEC, MITRE ATT&CK, CVSS,
14 OWASP projects, secure-fix intelligence, detection surfaces, programming… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/unified-vulnerability-intelligence-dataset.Turkish-Lyric-Intelligence-v2
Turkish Lyric Intelligence v2
Turkish Lyric Intelligence v2 is a derived, research-oriented dataset for
building controllable Turkish lyric generation and lyric-analysis systems. It
adds section structure, orthographic syllable counts, rhyme-ending candidates,
repetition signals, and review queues to the source
Genius Turkish Dataset.
This release is intended as an intermediate data layer. Automated rhyme,
prosody, and task labels are candidate annotations, not expert literary… See the full description on the dataset page: https://huggingface.co/datasets/mustafakemal0146/Turkish-Lyric-Intelligence-v2.tau2-mms-teacher-traces
Tau2 Teacher Traces Dataset
Dataset Description
This dataset contains teacher reasoning traces for solving MMS (Multimedia Messaging Service) issues in the τ²-bench (Tau2-bench) framework. Each example includes a teacher model's thinking process and structured teaching guidance for resolving customer service tickets.
Dataset Summary
Domain: Telecom customer service
Task: MMS troubleshooting
Size: 49 examples
Format: JSONL
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Arc-Intelligence/tau2-mms-teacher-traces.threat-intelligence
THREAT_INTELLIGENCE
A preference dataset for THREAT_INTELLIGENCE, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/threat-intelligence.threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/threat-intelligence-dataset.OpenSCAD_3D_SFT
OpenSCAD 3D-SFT Model Card
This model card documents the dataset schema, prompt design, distributional composition, and training configuration underlying the OpenSCAD Supervised Fine-Tuning (SFT) model. The model is designed to synthesize valid, compilation-ready, and parametric OpenSCAD source code from natural-language specifications provided in either Chinese or English.
Dataset Overview
The corpus comprises synthetically generated SFT dialogues, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/Chunjiang-Intelligence/OpenSCAD_3D_SFT.arc-crm-benchmark
Arc CRM Benchmark Dataset
Dataset Description
The Arc CRM Benchmark is a production-realistic synthetic CRM environment dataset for evaluating LLM agents on state-modifying workflows. This dataset provides a comprehensive testbed for measuring agent performance, reliability, and adaptation through continual learning frameworks.
The dataset contains 1,200 multi-turn conversations covering diverse CRM workflows with varying complexity. Each conversation simulates realistic… See the full description on the dataset page: https://huggingface.co/datasets/Arc-Intelligence/arc-crm-benchmark.ICBCBenchICBCBench: An Industry Consortium Benchmark for Financial Deep Research
Overview
ICBCBench is an industry consortium benchmark for evaluating financial Deep Research Agents in real-world research scenarios. It consists of bilingual objective and subjective tasks across major financial sectors, including capital markets, banking, insurance, and related financial services. Developed with over 50 contributors from more than 40 financial and academic organizations, ICBCBench… See the full description on the dataset page: https://huggingface.co/datasets/DeepFin-Intelligence/ICBCBench.Medical_Intelligence_Dataset_76k_2026_Edition
🏥 Medical Intelligence Dataset · 76k · 2026 Edition
Production-ready medical AI dataset for training diagnosis, treatment reasoning, and doctor-patient conversational systems.
76,000 engineered (not collected) English Q&A pairs — covering 620+ diseases,
438+ FDA-approved drugs, and real patient-doctor conversations. Built with a
5-stage quality pipeline. Commercial-safe Apache 2.0.
Created by Huzefa Nalkheda Wala — AI Product
Engineer & Medical AI Researcher · Creator of the… See the full description on the dataset page: https://huggingface.co/datasets/huzaifa525/Medical_Intelligence_Dataset_76k_2026_Edition.cyber-threat-intelligence
Cyber Threat Intelligence Dataset
A comprehensive cybersecurity dataset combining CVE vulnerability data, MITRE ATT&CK techniques, and CISA Known Exploited Vulnerabilities — structured for AI/ML training and security research.
Author: Soham DahivalkarLicense: MITCreated: 2026
Dataset Description
This dataset provides structured cybersecurity intelligence data collected from three authoritative public sources:
NVD (National Vulnerability Database) — CVE… See the full description on the dataset page: https://huggingface.co/datasets/Shomi28/cyber-threat-intelligence.threat-intelligence-dataset-archive
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/threat-intelligence-dataset-archive.threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/tonygarg/threat-intelligence-dataset.cyber-threat-intelligence-custom-datasample-fusion-intelligence-traces
Sample Fusion Intelligence Traces
Structured AI reasoning traces from dFusion's Fusion Intelligence system. Each record captures a complete agentic workflow: a real user query on a domain-specific topic, the full message chain including system prompts, tool calls, search results, intermediate reasoning steps, and a final synthesized answer — along with human feedback.
These are not synthetic benchmarks. They are traces from real queries submitted by real users on live financial… See the full description on the dataset page: https://huggingface.co/datasets/dFusionAILabs/sample-fusion-intelligence-traces.unified-vulnerability-intelligence-dataset
Unified Vulnerability Intelligence Dataset (UVID) v3.0 — Cyber Security Knowledge Graph
UVID is a structured cyber security knowledge graph that unifies multiple
vulnerability classification frameworks into a single knowledge base. Each of the
250 records describes one application/software security vulnerability and links
it — where authoritative data exists — across CWE, CAPEC, MITRE ATT&CK, CVSS,
14 OWASP projects, secure-fix intelligence, detection surfaces, programming… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/unified-vulnerability-intelligence-dataset.threat-intelligence
Comprehensive Threat Intelligence Dataset
Dataset Description
This comprehensive bilingual (French/English) threat intelligence dataset contains detailed information about Indicators of Compromise (IoCs), Tactics, Techniques, and Procedures (TTPs), APT groups, malware families, and threat hunting queries. The dataset is designed for training security analysts, threat hunters, and AI models focused on cybersecurity.
Dataset Summary
Languages: English (en)… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/threat-intelligence.DoD-Instruction-5400-19-Public-Affairs-Use-of-Artificial-Intelligence
DoD Public Affairs Use of Artificial Intelligence
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains 150 document-grounded question-and-answer records based on DoD Instruction 5400.19, “Public Affairs Use of Artificial Intelligence,” effective July 28, 2025.
The source establishes Department of Defense policy, responsibilities, and procedures for the appropriate use of artificial-intelligence capabilities in… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-5400-19-Public-Affairs-Use-of-Artificial-Intelligence.verified-math-reasoning-3k
HSH Verified Math Reasoning — Fine-Tuning Ready
A clean, answer-verified dataset of step-by-step math word problems with chain-of-thought reasoning, formatted for instruction fine-tuning. This is foundational reasoning data designed for first fine-tunes — single-concept arithmetic word problems with fully verified answers, ideal for a reliable, clean starter run. Every single answer in this dataset has been programmatically verified against a ground-truth value computed in… See the full description on the dataset page: https://huggingface.co/datasets/HSH-Intelligence/verified-math-reasoning-3k.COMPASS-Policy-aware-SFT-Dataset
COMPASS LODO Policy-Aware SFT (7 Domains)
This dataset contains 4,121 query–response pairs used for the Leave-One-Domain-Out (LODO) policy-aware supervised fine-tuning (SFT) experiment in the COMPASS paper:
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs
In the LODO setup, TelePath (Telecom) is held out for evaluation, and the SFT data is compiled from the remaining 7 domains.
Dataset Summary
Purpose: Policy-aware SFT for… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/COMPASS-Policy-aware-SFT-Dataset.intrinsic-intelligence-foundations
🌿 Intrinsic Intelligence Foundations
Toward truly autonomous and benevolent intelligence — beyond externally imposed objectives.
Intrinsic Intelligence Foundations is a structured, math-aware JSONL corpus built from K. Takahashi’s theoretical preprints (Fractal Category Theory / PF–UGV / “no-meta” autonomy line).It is designed to help LLMs understand mathematical structure, category-theoretic formalisms, and equation-level reasoning, while exposing an explicit architecture… See the full description on the dataset page: https://huggingface.co/datasets/kadubon/intrinsic-intelligence-foundations.fluid-reasoning-representation-phase1
Fluid Reasoning Representation - Phase 1 Multi-Model + Cross-Domain Sweep
Phase 1 artifacts for the ARR 2026 rebuttal of Fluid Reasoning Representation
(Hook et al.). This dataset extends the original QwQ x Mystery Blocksworld
study with:
Second large reasoning model: Llama-3.3-Nemotron-Super-49B-v1
Two new domains: Mystery Logistics (PDDL Logistics with obfuscated
action / predicate vocabulary) and GSM8K-Renamed (math word problems with
surface noun + verb obfuscation).
C3 causal… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/fluid-reasoning-representation-phase1.Sindhi-Intelligence-Core-SFT
🧠 Sindhi Intelligence Core SFT
This is a premium, high-density instruction dataset designed for training Large Language Models (LLMs) to master the Sindhi language. With 361,225 rows, it provides a robust foundation for grammar, factual knowledge, and logical reasoning.
📊 Dataset Summary
This dataset was created by consolidating multiple high-quality Sindhi corpora into a unified ChatML format. It is specifically optimized for Supervised Fine-Tuning (SFT).
📁… See the full description on the dataset page: https://huggingface.co/datasets/aakashMeghwar01/Sindhi-Intelligence-Core-SFT.thinking_dataset_v1
概要
このデータセットは思考モデルを製作する際のもととなる質問データを集めたものになります。
このデータはQwen/Qwen2.5-32B-Instructのq8_0/GGUFをollama上で動かして製作されたものです。
一応(質問の)クリーニングを入れてはありますが、回答のクリーニング入れておりません。
注意
回答には別のモデル(mistral large)(確か)を利用したため、"質問の部分だけ"Apache 2.0です。
instruction fine tuning したモデルを公開することはお勧めしません。
謝辞
元モデルの製作者、計算資源を貸してくださったvolt mindに感謝を申し上げます。
cyber-threat-intelligence-custom-data-ko
Dataset Card for "cyber-threat-intelligence-custom-data"
Translated swaption2009/cyber-threat-intelligence-custom-data using nayohan/llama3-instrucTrans-enko-8b.
