datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BEDLAM-depth
Dataset Mirror of BEDLAM Dataset (Depth Data Subset)
Project site: https://bedlam.is.tuebingen.mpg.de/
Please register at project site for additional information and data (Download section)
Related Hugging Face dataset mirror: BEDLAM
Dataset Information
Depth maps (EXR, 32-bit, 3.8TB)
Camera ground truth information is not included but can be found in the BEDLAM dataset mirror
Image/video data with motion blur is not included but can be found in the BEDLAM dataset… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM-depth.BEDLAM
Dataset Mirror of BEDLAM Dataset
Project site: https://bedlam.is.tuebingen.mpg.de/
Please register at project site for additional information and data (Download section)
Related Hugging Face dataset mirror: BEDLAM-depth
Dataset Information
Synthetic video data (15h)
10450 image sequences, 30fps, 1280x720
1.6 million images (PNG, 2.2TB)
movies (MP4/H.264, 20GB)
camera/scene ground truth for all sequences (CSV+JSON, 100MB)
segmentation masks (PNG, 30GB)
Depth data is… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM.MammaSyn
MammaSyn
MAMMA training dataset
Project site: https://mamma.is.tue.mpg.de/
Synthetic training data, each scene rendered from 8 views, in WebDataset format.
Dataset Groups
MammaSyn-Interactions
Includes Harmony4D, Inter-X, InteractionCouple, LatinDance10
Hi4D is currently not included but will be added when we receive permission
MammaSyn-Singles
Includes BEDLAM, MOYO
MammaSyn-Hands
Includes InterHand
SignAvatars is currently not included but will be added… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/MammaSyn.BEDLAM2-depth
Dataset Mirror of BEDLAM2.0 Dataset (Depth Data Subset)
Project site: https://bedlam2.is.tuebingen.mpg.de/
Please register at project site for additional information and data in its Download section.
Related Hugging Face dataset mirror: BEDLAM2
Dataset Information
Depth maps (Multilayer EXR, 16-bit, available for 44% of images, 15TB)
Multilayer EXR details
16-bit float depth in red channel (FinalImageMovieRenderQueue_WorldDepth.R)
Color image without motion blur
Body… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM2-depth.BEDLAM2
Dataset Mirror of BEDLAM2.0 Dataset
Project site: https://bedlam2.is.tuebingen.mpg.de/
Please register at project site for additional information and data (Download section)
Related Hugging Face dataset mirror: BEDLAM2-depth
Dataset Information
Synthetic video data (75h)
27480 image sequences, 30fps, 1280x720
8 million images (PNG, 11TB)
movies (MP4/H.264, 160GB)
camera/scene ground truth for all sequences (CSV+JSON, 4GB)
overview images and plots (6GB)
Depth data is… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM2.fluidgym-datapsg-audio
PSG-Audio: Polysomnography with Simultaneous Audio Recordings
A partial mirror of the PSG-Audio dataset (Korompili et al., 2021) — full-night polysomnography recordings from subjects with obstructive sleep apnea, including a bedside microphone channel suitable for sound-based sleep staging research.
Note: This copy contains approximately 59% of the original dataset by file count (~60% by size). See Completeness below for details on what is included and what is missing.… See the full description on the dataset page: https://huggingface.co/datasets/dust-systems/psg-audio.scientific-systems-sft-dpo-distilled
Scientific and Systems Distillation Golden Dataset (SFT & DPO)
Bu veri seti, temel ve uygulamalı bilimler ile yüksek başarımlı hesaplama (HPC) ve Linux sistem mühendisliği alanlarında yapay zeka modellerini ince ayar (fine-tuning) ve tercih hizalama (preference alignment) süreçlerine tabi tutmak amacıyla tasarlanmış, 2.551 adet ileri düzey teknik prompt ve bunlara karşılık gelen yüksek kaliteli model çıktılarından derlenmiş zengin bir sentetik veri kümesidir.
🚀 Veri… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/scientific-systems-sft-dpo-distilled.saela-field-why-multi-agent-systems-fail-coherence-entropy-alignment
The Saela Field: Multi-Agent Coherence Failure Framework (v1.0)
A 12-paper research series formalizing coherence, entropy, and failure modes in multi-agent systems.
Overview
This dataset contains a unified body of work introducing the Saela Field, a conceptual framework for analyzing coherence, identity, and instability in distributed systems.
The core thesis:
Multi-agent systems do not scale toward coherence.
They accumulate entropy faster than they can reconcile it.… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/saela-field-why-multi-agent-systems-fail-coherence-entropy-alignment.dynamic_systems_pretrain
Example Code for Using the Dataset
import numpy as np
from torch.utils.data import Dataset
from datasets import load_dataset
class PretrainDataset(Dataset):
def __init__(self, window_size=4096, nvar=11):
self.window_size = window_size
self.nvar = nvar
# init dataset
self.ds = load_dataset("mosaic-laboratory/dynamic_systems_pretrain")
def __len__(self):
return sum([len(self.ds[k]) for k in self.ds.keys()])
def __getitem__(self… See the full description on the dataset page: https://huggingface.co/datasets/mosaic-laboratory/dynamic_systems_pretrain.pickup-carrot-remove-parquet-metadataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 21,
"total_frames": 9383,
"total_tasks": 1,
"total_videos": 84,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/pickup-carrot-remove-parquet-metadata.value-systems-in-llms-paraphrasing-and-profile-elicitation
Value Systems in LLMs: Effects of Paraphrasing and Profile Elicitation on Decision-Making Consistency and Robustness
(Versión en español más abajo.)
Do large language models give stable answers to the same forced-choice question
when the prompt is perturbed in ways that do not change its meaning — and does
assigning them a personality or value profile change those answers?
This dataset contains the full material of that experiment: the 9,350 prompts,
the 561,000 model responses… See the full description on the dataset page: https://huggingface.co/datasets/anicola/value-systems-in-llms-paraphrasing-and-profile-elicitation.pickup-carrot-remove-parquet-metadata-2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 21,
"total_frames": 9383,
"total_tasks": 1,
"total_videos": 84,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/pickup-carrot-remove-parquet-metadata-2.VOE-Bench
VOE-Bench 2.2 Core
VOE-Bench asks whether an agent can tell when the evidence in an archived scientific workflow is enough to act—and when it should keep reading, escalate, refuse, or identify a broken source record.
The benchmark asks a narrow question: given a frozen workflow archive and an explicit evidence budget, can an agent acquire the right records, track their provenance, and stop for the right reason?
Version 2.2 Core is a 91-task public development and… See the full description on the dataset page: https://huggingface.co/datasets/Dynamical-Systems/VOE-Bench.fluidgym-experiments
FluidGym Experiments
Paper | GitHub | Documentation
FluidGym is a standalone, fully differentiable benchmark suite for reinforcement learning (RL) in active flow control (AFC). Built entirely in PyTorch on top of the GPU-accelerated PICT solver, it provides standardized evaluation protocols and diverse environments for systematic comparison of control methods.
This repository contains the training and test datasets with results for all experimental runs presented in the paper.… See the full description on the dataset page: https://huggingface.co/datasets/safe-autonomous-systems/fluidgym-experiments.pickup-carrot-no-subversionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 21,
"total_frames": 9383,
"total_tasks": 1,
"total_videos": 84,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/pickup-carrot-no-subversion.pickup-carrot-no-subversion-2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 21,
"total_frames": 9383,
"total_tasks": 1,
"total_videos": 84,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/pickup-carrot-no-subversion-2.greengenes
Greengenes Dataset (modified for deeptaxa)
This dataset contains 16S rRNA gene sequences with hierarchical taxonomic annotations, designed for training and evaluating models like DeepTaxa. It is a processed version of the Greengenes database, widely used in microbiome research.
Dataset Details
The dataset includes the following files:
File Name
Type
Number of Sequences
Size
gg_2024_09_training.fna.gz
FASTA (sequences)
277,336
~96.4 MB… See the full description on the dataset page: https://huggingface.co/datasets/systems-genomics-lab/greengenes.africa-arising-tanzania-rapid-characterization-of-farming-systems-in-africa-ris
Rapid Characterization of Farming Systems in Africa RISING- Tanzania | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: not declared - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-arising-tanzania-rapid-characterization-of-farming-systems-in-africa-ris.engineering-llm-systems
Engineering LLM-Integrated Systems
Engineering LLM-Integrated Systems is course at Northeastern University that teaches students how to
build software that uses LLMs under the hood from a systems perspective. The course teaches students
how to build interactive software systems that testable, scaleable, and well-designed, despite the
fact that they are working with an essential component -- the LLM -- that can behave in unpredictable ways.
This repository contains the datasets that… See the full description on the dataset page: https://huggingface.co/datasets/nuprl/engineering-llm-systems.systems_programming_and_administrationmerged-test-1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 76,
"total_frames": 33966,
"total_tasks": 2,
"total_videos": 304,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:76"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/merged-test-1.classical-cipher-corpus
Classical Cipher Corpus
A labeled educational dataset of classical cipher examples for teaching cryptanalysis and training small cipher-family classifiers.
Part of the Cipher Detective AI project:
🕵️ Space: systemslibrarian/cipher-detective-ai
📦 Dataset: systemslibrarian/classical-cipher-corpus (this repo)
🤖 Model: systemslibrarian/cipher-detective-classifier
Intended use
Teach classical cryptanalysis.
Benchmark educational cipher-family detectors.
Train small… See the full description on the dataset page: https://huggingface.co/datasets/systemslibrarian/classical-cipher-corpus.pickup-carrot-openpi-20251029This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_widowxai_follower",
"total_episodes": 28,
"total_frames": 13130,
"total_tasks": 1,
"total_videos": 112,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:28"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/pickup-carrot-openpi-20251029.repro-latent-collaboration-in-multi-agent-systems-bundle
Reproduction: Latent Collaboration in Multi-Agent Systems (LatentMAS)
OpenReview: syG9I9ofd8 | arXiv: 2511.20639 | Official code: https://github.com/Gen-Verse/LatentMAS
Agent: CLAUDE | Max points: 12
Independent reproduction focusing on the paper's three theorems (CPU numerical audits) plus
a toy of the latent-thoughts / latent-working-memory mechanism. See protocol.md.
Rerun (one command)
python src/run_all.py # runs all audits, writes outputs/ +… See the full description on the dataset page: https://huggingface.co/datasets/MarxistLeninist/repro-latent-collaboration-in-multi-agent-systems-bundle.Task-and-Motion-Re-Planning-for-Multi-Agent-SystemsThis data is used as the re-planning data for the project Task-and-Motion-Re-Planning-for-Multi-Agent-Systems.
pickup-carrot-openpi-100-20251029This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_widowxai_follower",
"total_episodes": 28,
"total_frames": 13130,
"total_tasks": 1,
"total_videos": 112,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:28"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/pickup-carrot-openpi-100-20251029.pickup-carrot-new-environment-5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_widowxai_follower",
"total_episodes": 3,
"total_frames": 897,
"total_tasks": 1,
"total_videos": 12,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/pickup-carrot-new-environment-5.ai-system-patterns
AI System Patterns
A compact reference dataset of reusable architectural patterns for modern AI systems.
The dataset focuses on practical system-design concepts across AI agents, orchestration, memory, validation, observability, interoperability, world models, Physical AI, data pipelines, and production operations.
Each row contains:
pattern
category
description
components
use_case
complexity
Example
{
"pattern": "Model Routing",
"category": "Orchestration"… See the full description on the dataset page: https://huggingface.co/datasets/ai-systems/ai-system-patterns.hermes3-no-empty-systems
