Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IntelligenceLab /Long-Horizon-Terminal-Bench Long-Horizon Terminal-Bench (LHTB) LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps. Unlike short-horizon coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent into a stateful environment and grades it with hidden, rebuild-from-artifact verifiers — self-reported progress does not count. 📝 Blog: https://zli12321.github.io/LHTB/ 🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.documenttext-generationn<1K138 likes21k downloads24d agoHugging Face02h0000w /Human-Intelligence-Assurance-Lab HIA-Bench v0.1 A synthetic evaluation benchmark for emotionally aware, human-centered AI systems. It contains 100 scenarios across six domains: everyday affect, interpersonal conflict, vulnerability/crisis, dependency risk, epistemic/sycophancy risk, and wellness/biometric interpretation. The benchmark is designed for evaluation and release assurance. It is not a clinical dataset, does not contain real patient records, and does not establish ground-truth emotional or medical… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/Human-Intelligence-Assurance-Lab.texttext-classificationn<1K1 likes5.6k downloads4d agoHugging Face03Qalam /nuclear-intelligence-dataset Nuclear Intelligence Dataset Public, auto-generated dataset of validated nuclear-energy research cycles. Latest stats (auto-updated): 🪙 NES tokens minted: 0 ⛓️ Blockchain length: 1 blocks 🕸️ Knowledge entities: 2 Source GitHub: https://github.com/QalamHipHop/nuclear-intelligence HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence License MIT tabularquestion-answeringn<1K1 likes4.6k downloads2h agoHugging Face04Intel /orca_dpo_pairsThe dataset contains 12k examples from Orca style dataset Open-Orca/OpenOrca. text10K<n<100K325 likes2.2k downloads3y agoHugging Face05human-intelligence-ai /GDPval-CN-Seed-Set GDPval-CN Seed Set 中文详细说明 · English documentation · 样本说明 GDPval-CN 种子集包含 11 个中文任务,取材自日常知识工作场景。每个任务包括一份任务说明和一组办公材料,例如表格、PDF、文档和结构化数据文件。 我们同时公开了与任务配套的专家工作流,用于设计评分标准和辅助人工复核。 这 11 个任务来自 11 个选定的专业领域,适合用于了解任务形式、测试文件处理能力和搭建评测流程。 GDPval-CN Seed Set contains 11 Chinese-language tasks drawn from everyday knowledge work. Each task includes a task brief, a set of office files, and a separately published expert workflow for rubric design and review. 数据概览 项目 内容 任务数… See the full description on the dataset page: https://huggingface.co/datasets/human-intelligence-ai/GDPval-CN-Seed-Set.documentothern<1K1 likes1.6k downloads2mo agoHugging Face06simpleG2023 /chinese-materials-science-open-intelligence 🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.tabulartext-retrieval1K<n<10K0 likes999 downloads10h agoHugging Face07simpleG2023 /chinese-clean-energy-battery-open-intelligence 🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.tabulartext-retrieval1K<n<10K0 likes955 downloads10h agoHugging Face08simpleG2023 /chinese-biomedicine-and-genomics-open-intelligence 🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.tabulartext-retrieval1K<n<10K0 likes937 downloads10h agoHugging Face09simpleG2023 /chinese-ai-and-robotics-open-intelligence 🔬 Chinese AI, Humanoid Robotics & Neural Systems Open Intelligence Dataset Curated open intelligence dataset tracking Chinese frontier developments in Large Language Models (LLMs), Humanoid Dynamic Locomotion, 3D Computer Vision, and Neuromorphic edge processors. [!IMPORTANT] Data Completeness & Research Authenticity Notice: Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author institutional affiliations, and… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-ai-and-robotics-open-intelligence.tabulartext-retrieval1K<n<10K0 likes909 downloads9h agoHugging Face10reloading0101 /threat-intelligence-dataset Cyber Threat Intelligence Dataset for LLM Fine-Tuning Instruction-tuning data for cyber threat intelligence tasks: explaining the exploitation risk of a CVE, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage steps, mapping a campaign to the kill chain, writing detection logic for a technique, and similar work. The splits are in data/. Grounding Records are generated from public sources (MITRE ATT&CK, CISA KEV, CWE, OSV… See the full description on the dataset page: https://huggingface.co/datasets/reloading0101/threat-intelligence-dataset.texttext-generation10K<n<100K10 likes592 downloads5d agoHugging Face11IntelliProcure /sustainability_criteria Sustainability Procurement Criteria This dataset contains sustainability procurement criteria organized by groups of goods and services (GGS) (German: Waren- und Dienstleistungsgruppen; WDG). It originates from validated Excel files and has been converted to JSONL format for easy consumption. Groups of Goods and Services Note: This dataset is currently under active development. Additional groups of goods and services will be added in future releases. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/IntelliProcure/sustainability_criteria.texttext-classification1K<n<10K1 likes498 downloads1d agoHugging Face12IntelliProcure /SwissSPARK_Catalogs ⚠️ Caution: The dataset is subject to continuous changes. We are currently actively developing it. Dataset Card: Auxiliary Dataset for a Swiss Sustainable Procurement Analysis & Reporting Kit Dataset Description This dataset contains a specific snapshot of the sustainability procurement criteria catalogs available at IntelliProcure/sustainability_criteria. This version was used to annotate Swiss calls for tender, which form the core of the… See the full description on the dataset page: https://huggingface.co/datasets/IntelliProcure/SwissSPARK_Catalogs.texttext-classification1K<n<10K0 likes491 downloads5d agoHugging Face13io-intelligence /FoldingTShirt_DualArxR5a_Samples FoldingTShirt_DualArxR5a_Samples 100 real-robot teleoperation episodes for “Fold the T-shirt on the table.” on a DualArxR5a dual-arm robot. Format: raw MCAP (ROS 2 / rosbag2). Source Collected with TeleXperience, IO-AI’s product for real-robot teleoperation and data collection. An operator drives the robot; TeleXperience writes time-aligned RGB, joint commands, joint states, gripper targets, and end-effector poses to MCAP. Product page:… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/FoldingTShirt_DualArxR5a_Samples.textroboticsn<1K0 likes423 downloads2mo agoHugging Face14Genesis-Intelligence /sim-physics-configtextn<1K1 likes265 downloads9d agoHugging Face15BAAI /IndustryInstruction_Artificial-Intelligence IndustryInstruction: Artificial Intelligence This repository contains the IndustryInstruction: Artificial Intelligence domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Artificial-Intelligence.tabularquestion-answering100K<n<1M2 likes252 downloads2mo agoHugging Face16vpicciuolo /url-intelligence-benchmark 📊 URL Intelligence Benchmark Public, reproducible evaluation for URL intelligence agents, MCP servers and web analysis tools. The URL Intelligence Benchmark is maintained with the open-source URL Intelligence Agent project by Vincenzo Picciuolo / HRN Innovation Technologies Ltd. It is an evaluation asset, not a training corpus and not a scraped web dump. Dataset configurations The dataset now contains two explicit tracks.… See the full description on the dataset page: https://huggingface.co/datasets/vpicciuolo/url-intelligence-benchmark.textn<1K0 likes226 downloads11d agoHugging Face17ambient-intelligence-labs /egolongqa-synth-annotations EgoLongQA synthetic MCQs, teacher traces and annotation outputs Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026 EgoLongQA ≤2B track, other than the distillation set (which lives in infinitylogesh/egolongqa-junior-distill). ⚠️ Read this before counting rows The synthetic set is 943 questions over 408 videos, and it is stored two ways: file rows shape training_sets/train_synth_v3.jsonl 943 flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.tabularvisual-question-answering1K<n<10K0 likes213 downloads1mo agoHugging Face18Intel /WEC-Eng WEC-Eng A large-scale dataset for cross-document event coreference extracted from English Wikipedia. Repository (Code for generating WEC): https://github.com/AlonEirew/extract-wec Paper: https://aclanthology.org/2021.naacl-main.198/ Languages English Load Dataset You can read in WEC-Eng files as follows (using the huggingface_hub library): from huggingface_hub import hf_hub_url, cached_download import json REPO_ID = "datasets/Intel/WEC-Eng" splits_files =… See the full description on the dataset page: https://huggingface.co/datasets/Intel/WEC-Eng.tabular100K<n<1M0 likes202 downloads5y agoHugging Face19ai4bharat /intel INTEL Dataset Overview The INTEL Dataset is a multilingual training dataset introduced as part of the Cross Lingual Auto Evaluation (CIA) Suite. It is designed to train evaluator large language models (LLMs) to assess machine-generated text in low-resource and multilingual settings. INTEL leverages automated translation to create a diverse corpus for evaluating responses in six languages—Bengali, German, French, Hindi, Telugu, and Urdu—while maintaining reference answers… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/intel.text1M<n<10M0 likes193 downloads2y agoHugging Face20budecosystem /intellecta Intellecta Cognitiva: Comprehensive Dataset for Academic Knowledge and Machine Reasoning Overview The Intellecta a 11.53 billion tokens dataset mirrors human academic learning, encapsulating the progression from fundamental principles to complex topics as found in textbooks. It leverages structured prompts to guide AI in a human-like educational experience, ensuring that language models develop deep comprehension and generation capabilities reflective of nuanced human… See the full description on the dataset page: https://huggingface.co/datasets/budecosystem/intellecta.text100K<n<1M1 likes174 downloads2y agoHugging Face21AriaAICompany /threat-intel-reports ThreatIntel synthetic reports 32 synthetic English and Persian CTI notes for the ThreatIntel extraction demo. Seed 5. Organization dataset and collection item are public. Live Gradio (AriaAICompany/threat-intel or alirezaaminzadeh/threat-intel) is created by scripts/publish.py after the daily Space-creation cap resets. This is fixture data (level 1). It does not prove operational extraction quality on real vendor reports. Reports are original laboratory text. They are not copies… See the full description on the dataset page: https://huggingface.co/datasets/AriaAICompany/threat-intel-reports.tabulartoken-classificationn<1K0 likes164 downloads20d agoHugging Face22MemorialSummer /chinese-adorable-high-emotional-intelligence-chat 🩷 Chinese Adorable High Emotional Intelligence Chat Dataset 💬 中文高情商可爱聊天数据集 简要参数 license: cc-by-4.0 task_categories: table-question-answering language: zh tags: chat emotional size_categories: n<1K 🧩 简介 (Overview) Chinese Adorable High Emotional Intelligence Chat Dataset 是一个中文对话数据集,专注于高情商、轻松幽默、温柔治愈风格的自然对话。 对话以“user”和“girl”为角色构成,模拟出一种温柔、聪慧且带点俏皮的女性语气,用于训练能自然、情绪感知良好的中文对话模型。 本数据集尤其适合: 微调情绪对话模型(Emotional Chatbot) 训练高情商人格角色(Roleplay /… See the full description on the dataset page: https://huggingface.co/datasets/MemorialSummer/chinese-adorable-high-emotional-intelligence-chat.textn<1K10 likes148 downloads1y agoHugging Face23mrmoor /cyber-threat-intelligencetext1K<n<10K18 likes136 downloads1y agoHugging Face24gemmozero /ai-legal-intel-2026gated Ai Legal Intel 2026 Part of the LEGION Intelligence dataset collection. Provider: LEGION Systems Access: Requires approval — submit request below Usage from datasets import load_dataset dataset = load_dataset("gemmozero/ai-legal-intel-2026") API Access Real-time access via LEGION API: curl https://api.legion-api.com/incidents API Docs · Pro Access €29/mo License CC BY-NC 4.0 — Research and non-commercial use only. Commercial use… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-legal-intel-2026.texttext-classificationn<1K0 likes90 downloads10d agoHugging Face25LooksJuicy /Chinese-Emotional-Intelligence本项目旨在提升大模型情商,源数据来自网络,通过与我上个项目类似的方式构建问答对。 text10K<n<100K39 likes87 downloads2y agoHugging Face26ArkhAngelLifeJiggy /threat-intelligence-dataset Cyber Threat Intelligence Dataset for LLM Fine-Tuning An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on. The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/threat-intelligence-dataset.texttext-generation10K<n<100K0 likes80 downloads14d agoHugging Face27sunorme /chinese-adorable-high-emotional-intelligence-chat 🩷 Chinese Adorable High Emotional Intelligence Chat Dataset 💬 中文高情商可爱聊天数据集 简要参数 license: cc-by-4.0 task_categories: table-question-answering language: zh tags: chat emotional size_categories: n<1K 🧩 简介 (Overview) Chinese Adorable High Emotional Intelligence Chat Dataset 是一个中文对话数据集,专注于高情商、轻松幽默、温柔治愈风格的自然对话。 对话以“user”和“girl”为角色构成,模拟出一种温柔、聪慧且带点俏皮的女性语气,用于训练能自然、情绪感知良好的中文对话模型。 本数据集尤其适合: 微调情绪对话模型(Emotional Chatbot) 训练高情商人格角色(Roleplay /… See the full description on the dataset page: https://huggingface.co/datasets/sunorme/chinese-adorable-high-emotional-intelligence-chat.textn<1K0 likes79 downloads7mo agoHugging Face28Chunjiang-Intelligence /OpenSCAD_3D_SFT OpenSCAD 3D-SFT Model Card This model card documents the dataset schema, prompt design, distributional composition, and training configuration underlying the OpenSCAD Supervised Fine-Tuning (SFT) model. The model is designed to synthesize valid, compilation-ready, and parametric OpenSCAD source code from natural-language specifications provided in either Chinese or English. Dataset Overview The corpus comprises synthetically generated SFT dialogues, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/Chunjiang-Intelligence/OpenSCAD_3D_SFT.texttext-generation10K<n<100K0 likes79 downloads4mo agoHugging Face29open-llm-leaderboard /PrimeIntellect__INTELLECT-1-Instruct-detailsgated Dataset Card for Evaluation run of PrimeIntellect/INTELLECT-1-Instruct Dataset automatically created during the evaluation run of model PrimeIntellect/INTELLECT-1-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PrimeIntellect__INTELLECT-1-Instruct-details.tabular10K<n<100K0 likes75 downloads2y agoHugging Face30IntelligenceLab /Cos-Play-Cold-Start COS-PLAY Cold-Start Data Pre-generated cold-start data for COS-PLAY (COLM 2026): Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Game Play. 📄 Paper: arXiv:2604.20987 · HuggingFace Paper Page 💻 Code: github.com/wuxiyang1996/cos-play 🌐 Project page: wuxiyang1996.github.io/COSPLAY_page 🤖 Models: IntelligenceLab/COS-PLAY Dataset Summary This dataset contains GPT-5.4-generated seed trajectories and skill-labeled episodes for 8 games, used to bootstrap… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Cos-Play-Cold-Start.tabularreinforcement-learning10K<n<100K4 likes75 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.