datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
awesome-loop-engineering
Awesome Loop Engineering Dataset
A structured dataset of 1025 papers, official docs, tools, benchmarks, patterns, critiques, and implementation guides for recurring AI-agent systems.
Resource Atlas ·
GitHub field guide ·
Resource selection ·
Report a correction
Dataset Summary
Each row connects an original source to its contribution, novelty, impact, publication details, lifecycle stages, audience, evidence type, link status, and… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-loop-engineering.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.CineBoard3D-plus
🎬 CineBoard3D++: Dynamic 3D Story World Dataset
📊 Dataset Summary
CineBoard3D++ is a collection of editable, movie-inspired 3D story worlds built with StoryBlender for narrative-grounded camera planning and world visual attention. It brings together story scripts, animated characters, scene geometry, and shot-level configurations in native Blender projects.
The benchmark covers 50 stories, 457 scenes, 1,585 shots, and 3,197 3D assets (836 plot-related and 2,361… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringAI-LAB/CineBoard3D-plus.PLSemanticsBench
The 43rd International Conference on Machine Learning (ICML 2026), Seoul, South Korea
LLMs Lean on Priors, Not Programming Language
Semantics
by
Aditya Thimmaiah1,
Jiyang Zhang1,
Jayanth Srinivasa2,
Junyi Jessy Li1,
Milos Gligoric1
1The University of Texas at Austin
2Cisco Research
TLDR: Frontier LLMs execute programs with up to 90–100% accuracy when symbols retain their usual… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringSoftware/PLSemanticsBench.reverse_engineering_code_dataset_O1_x86_O1awesome-graph-engineering
Awesome Graph Engineering Resource Atlas
A versioned collection of research, standards, frameworks, protocols, reliability systems, evaluations, and critiques for graph-structured multi-agent systems and programmable AI-agent organizations.
This dataset mirrors Awesome Graph Engineering. The GitHub JSONL file is canonical; the Hub exposes the same records through Dataset Viewer, direct downloads, datasets, and pandas.
Working definition
Graph engineering is the… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/awesome-graph-engineering.arxiv-software-engineering-datasetsoftware-engineering-interview-practices-2005-2026
Replication Package: Yesterday's Interviews for Today's Engineers
This repository contains the de-identified analytical data and the Python
reproduction script for:
Vitalii Romaniuk. "Yesterday's Interviews for Today's Engineers:
Retrospective Perceptions and a Work-Aligned Hiring Framework (2005–2026)."
arXiv:2609.14046, 2026.
Paper: https://arxiv.org/abs/2609.14046
Contents
data/survey_responses_deidentified.csv contains the 911 retained survey
records used… See the full description on the dataset page: https://huggingface.co/datasets/Kidomakai/software-engineering-interview-practices-2005-2026.engineering-llm-systems
Engineering LLM-Integrated Systems
Engineering LLM-Integrated Systems is course at Northeastern University that teaches students how to
build software that uses LLMs under the hood from a systems perspective. The course teaches students
how to build interactive software systems that testable, scaleable, and well-designed, despite the
fact that they are working with an essential component -- the LLM -- that can behave in unpredictable ways.
This repository contains the datasets that… See the full description on the dataset page: https://huggingface.co/datasets/nuprl/engineering-llm-systems.llm-engineering-journey-datacomplete-2026-ai-engineering-master-suite
👑 Complete 2026 AI Engineering & Research Master Suite (10-in-1 Mega-Bundle)
The definitive all-access research intelligence repository featuring 13,246 domain-verified research papers and 2,500+ official code repositories spanning all 10 core AI domains: Code Generation (SWE-bench), Real-Time Voice Agents (<200ms), Robotics & Embodied AI, Multi-Agent Swarms, Enterprise RAG, Medical AI, Multimodal Vision-Language, AI Security & Red-Teaming, Crypto AI Agents, and LLM… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/complete-2026-ai-engineering-master-suite.APED-African-Protein-Engineering-Dataset
APED African Protein Engineering Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: parquet - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/APED-African-Protein-Engineering-Dataset.cmu-engineering-publicationsreverse_engineering_code_dataset_O0_arm_O0reverse_engineering_code_dataset_O2_x64_O2Mechanical-engineering
Dataset Overview
This dataset contains prompt–response pairs generated using OpenAI’s GPT-4o-mini model via the Azure OpenAI API.
The dataset was created for fine-tuning and research on mechanical engineering Q&A.
Model: GPT-4o-mini
License: CC-BY-4.0
Date: November 2025
Notes: All responses are synthetic; no private or proprietary data included.
reverse_engineering_code_dataset_O2_arm_O2reverse_engineering_code_dataset_O3_arm_O3reverse_engineering_code_dataset_O3_x64_O3reverse_engineering_code_dataset_O2_mips_O2reverse_engineering_code_dataset_O1_arm_O1reverse_engineering_code_dataset_O0_x64_O0reverse_engineering_code_dataset_O1_mips_O1reverse_engineering_code_dataset_O1_x64_O1reverse_engineering_code_dataset_O0_mips_O0MMLU-Pro-engineering-outputs
Load the Dataset
from datasets import load_dataset
ds = load_dataset("eewer/MMLU-Pro-engineering-outputs", split="train")
Dataset Features
Feature Name
Description
question_id
(int64) Unique identifier for each question in the dataset.
question
(string) The text of the question being asked.
options
(list of string) Multiple-choice answer options for the question.
answer
(string) The correct answer for the question, as a text string.
answer_index… See the full description on the dataset page: https://huggingface.co/datasets/eewer/MMLU-Pro-engineering-outputs.recreated_reverse_engineering_code_dataset_O1_x86_O1recreated_reverse_engineering_code_dataset_O3_x86_O3reverse_engineering_code_dataset_O3_x86_O3context-engineering-v1
Context Engineering V1: Sequential API Recommendation Dataset
This dataset accompanies the research paper:
Rethink Context Engineering Using an Attention-based Architecture
Yiqiao Yin — University of Chicago Booth School of Business / Columbia University
It was generated using the open-source context-engineer Python package:
GitHub: https://github.com/yiqiao-yin/context-engineer-repo
PyPI: https://pypi.org/project/context-engineer/0.1.0/
Dataset Summary
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/context-engineering-v1.
