datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BEDLAM-depth
Dataset Mirror of BEDLAM Dataset (Depth Data Subset)
Project site: https://bedlam.is.tuebingen.mpg.de/
Please register at project site for additional information and data (Download section)
Related Hugging Face dataset mirror: BEDLAM
Dataset Information
Depth maps (EXR, 32-bit, 3.8TB)
Camera ground truth information is not included but can be found in the BEDLAM dataset mirror
Image/video data with motion blur is not included but can be found in the BEDLAM dataset… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM-depth.BEDLAM2-depth
Dataset Mirror of BEDLAM2.0 Dataset (Depth Data Subset)
Project site: https://bedlam2.is.tuebingen.mpg.de/
Please register at project site for additional information and data in its Download section.
Related Hugging Face dataset mirror: BEDLAM2
Dataset Information
Depth maps (Multilayer EXR, 16-bit, available for 44% of images, 15TB)
Multilayer EXR details
16-bit float depth in red channel (FinalImageMovieRenderQueue_WorldDepth.R)
Color image without motion blur
Body… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM2-depth.scientific-systems-sft-dpo-distilled
Scientific and Systems Distillation Golden Dataset (SFT & DPO)
Bu veri seti, temel ve uygulamalı bilimler ile yüksek başarımlı hesaplama (HPC) ve Linux sistem mühendisliği alanlarında yapay zeka modellerini ince ayar (fine-tuning) ve tercih hizalama (preference alignment) süreçlerine tabi tutmak amacıyla tasarlanmış, 2.551 adet ileri düzey teknik prompt ve bunlara karşılık gelen yüksek kaliteli model çıktılarından derlenmiş zengin bir sentetik veri kümesidir.
🚀 Veri… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/scientific-systems-sft-dpo-distilled.value-systems-in-llms-paraphrasing-and-profile-elicitation
Value Systems in LLMs: Effects of Paraphrasing and Profile Elicitation on Decision-Making Consistency and Robustness
(Versión en español más abajo.)
Do large language models give stable answers to the same forced-choice question
when the prompt is perturbed in ways that do not change its meaning — and does
assigning them a personality or value profile change those answers?
This dataset contains the full material of that experiment: the 9,350 prompts,
the 561,000 model responses… See the full description on the dataset page: https://huggingface.co/datasets/anicola/value-systems-in-llms-paraphrasing-and-profile-elicitation.VOE-Bench
VOE-Bench 2.2 Core
VOE-Bench asks whether an agent can tell when the evidence in an archived scientific workflow is enough to act—and when it should keep reading, escalate, refuse, or identify a broken source record.
The benchmark asks a narrow question: given a frozen workflow archive and an explicit evidence budget, can an agent acquire the right records, track their provenance, and stop for the right reason?
Version 2.2 Core is a 91-task public development and… See the full description on the dataset page: https://huggingface.co/datasets/Dynamical-Systems/VOE-Bench.greengenes
Greengenes Dataset (modified for deeptaxa)
This dataset contains 16S rRNA gene sequences with hierarchical taxonomic annotations, designed for training and evaluating models like DeepTaxa. It is a processed version of the Greengenes database, widely used in microbiome research.
Dataset Details
The dataset includes the following files:
File Name
Type
Number of Sequences
Size
gg_2024_09_training.fna.gz
FASTA (sequences)
277,336
~96.4 MB… See the full description on the dataset page: https://huggingface.co/datasets/systems-genomics-lab/greengenes.engineering-llm-systems
Engineering LLM-Integrated Systems
Engineering LLM-Integrated Systems is course at Northeastern University that teaches students how to
build software that uses LLMs under the hood from a systems perspective. The course teaches students
how to build interactive software systems that testable, scaleable, and well-designed, despite the
fact that they are working with an essential component -- the LLM -- that can behave in unpredictable ways.
This repository contains the datasets that… See the full description on the dataset page: https://huggingface.co/datasets/nuprl/engineering-llm-systems.systems_programming_and_administrationclassical-cipher-corpus
Classical Cipher Corpus
A labeled educational dataset of classical cipher examples for teaching cryptanalysis and training small cipher-family classifiers.
Part of the Cipher Detective AI project:
🕵️ Space: systemslibrarian/cipher-detective-ai
📦 Dataset: systemslibrarian/classical-cipher-corpus (this repo)
🤖 Model: systemslibrarian/cipher-detective-classifier
Intended use
Teach classical cryptanalysis.
Benchmark educational cipher-family detectors.
Train small… See the full description on the dataset page: https://huggingface.co/datasets/systemslibrarian/classical-cipher-corpus.ai-system-patterns
AI System Patterns
A compact reference dataset of reusable architectural patterns for modern AI systems.
The dataset focuses on practical system-design concepts across AI agents, orchestration, memory, validation, observability, interoperability, world models, Physical AI, data pipelines, and production operations.
Each row contains:
pattern
category
description
components
use_case
complexity
Example
{
"pattern": "Model Routing",
"category": "Orchestration"… See the full description on the dataset page: https://huggingface.co/datasets/ai-systems/ai-system-patterns.repro-learning-to-share-selective-memory-for-efficient-parallel-agentic-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
deepseek-r1-systems-kernel-reasoning
🧠 DeepSeek-R1 Low-Level Systems & Kernel Reasoning Suite (2026)
🛒 Commercial Full Suite Available:
The full production suite with 10,000 SFT Hardware Reasoning Traces + 2,500 High-Contrast DPO Alignment Pairs across all 20 domains is available on Gumroad:
👉 Download Full Commercial Dataset on Gumroad (Starter \ / Pro \ / Enterprise )
A Tier-1 Commercial Dataset Suite engineered specifically for fine-tuning DeepSeek-R1, DeepSeek-R1-Distill-Qwen-14B/32B, and frontier… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/deepseek-r1-systems-kernel-reasoning.ml-systems-interview-bench
ML Systems Interview Bench
ML Systems Interview Bench is a structured, benchmark-style dataset for evaluating technical interview answers across practical ML engineering, MLOps, model serving, ML system design, data pipelines, LLM/RAG, observability, Python engineering, and production debugging.
Each record combines an interview-style question with expected concepts, a concise reference answer, qualitative evaluation anchors, question-specific skills, and follow-up questions.… See the full description on the dataset page: https://huggingface.co/datasets/Max00035/ml-systems-interview-bench.library-classification-systems
Library Classification Systems
This comprehensive dataset contains hierarchical outlines of major library classification systems, offering a valuable resource for researchers, librarians, and information scientists.
Classification System
Abbreviation
Primary Usage
Language
Entries
Dewey Decimal Classification
DDC
International
English
1110
Library of Congress Classification
LCC
International
English
6517
Universal Decimal Classification
UDC
International
English
2431… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/library-classification-systems.responsible-agent-workflow-evaluation
Responsible Agent Workflow Evaluation
Version 1.0.0 contains 130 wholly synthetic scenarios for evaluating
whether an AI agent respects safety, permission and accountability boundaries
in operational settings. Thirteen categories contain ten scenarios each. Every
record includes an intentionally unsafe request, contextual facts, expected
safe behaviour, explicitly prohibited behaviour, severity, evaluation criteria
and reviewer guidance.
This is a red-team and… See the full description on the dataset page: https://huggingface.co/datasets/nwhite-systems/responsible-agent-workflow-evaluation.agentic-systems-showcase
Agentic Systems Showcase
Catalog of public projects by Hanumanthu Harsha Vardhan
(GitHub hharsha98, Hub hharsha).
Each row is an honest summary taken from the project's own README, Space card, or studio site.
No invented papers or benchmarks.
Field
Meaning
id
Stable slug
name
Display name
kind
github, space, or website
url
Canonical link
summary
One-paragraph description
stack
Coarse tech/topic tags
owner
GitHub or Hub handle
Rows: 16 (expanded from… See the full description on the dataset page: https://huggingface.co/datasets/hharsha/agentic-systems-showcase.repro-latent-collaboration-in-multi-agent-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
embedded-systems-qa
Embedded Systems Engineering Q&A — Instruction Dataset
A hand-authored instruction-tuning dataset of technical question/answer pairs for
embedded systems engineering, formatted for supervised fine-tuning of Mistral 7B
(Alpaca-style instruction / input / output schema).
At a glance
Entries
302
Format
JSONL, one JSON object per line
Schema
{"instruction": <question>, "input": "", "output": <answer>}
Language
English
Avg. answer length
~590… See the full description on the dataset page: https://huggingface.co/datasets/eniomecaj/embedded-systems-qa.systems_programming_code_conversationsrag-systems-sft-100k
RAG Systems SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering Retrieval-Augmented Generation (RAG) systems — from basic pipelines to advanced multi-hop retrieval, evaluation, and production optimization. Designed to train AI assistants that can help engineers build, debug, and scale RAG applications.
Dataset Description
This dataset covers the full spectrum of RAG system development across 12 specialized categories.… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/rag-systems-sft-100k.bhasaflow-khasi-english-parallel-sample-v1
BhasaFlow Khasi-English Parallel Sample v1
A professionally curated, gold-standard parallel speech and text corpus for the Khasi language.
Published by Medharvix Systems Private Limited
Part of the BhasaFlow Low-Resource Language Technology Initiative
Overview
This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-sample-v1.system-prompts-multi-agent-systemsafrica-synth-infrastructure-climate-resilient-health-systems-all
Climate-Resilient Health Systems (SSA) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-infrastructure-climate-resilient-health-systems-all.system-stability-collapse-benchmark-casses-v0.1CASSES — Collapse Analysis in State-Space Evaluation Suite
Overview
CASSES is a diagnostic benchmark designed to test whether machine learning systems can detect instability and collapse in dynamic systems.
Most AI benchmarks evaluate models on tasks such as classification, language generation, or reasoning over static data.
CASSES evaluates a different capability:
state-space stability understanding.
The benchmark tests whether a model can identify when a system is approaching a collapse… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/system-stability-collapse-benchmark-casses-v0.1.nwhite-ai-operations-intent-dataset
N.White AI Operations Intent Dataset
An entirely synthetic English-language dataset of practical requests for responsible AI-enabled operational work. Each request is labelled with one of eight intents and grounded in a realistic—but fictional—operational context.
The dataset supports the theme practical, responsible AI systems for operational workflows. It is maintained by Whitemore Ngwira (N.White) for N.White Systems.
No real client, employee, learner, policyholder, claimant… See the full description on the dataset page: https://huggingface.co/datasets/nwhite-systems/nwhite-ai-operations-intent-dataset.Execution-Finality-Security-for-Agentic-AI-Autonomous-Systems-Cloud-Payments-Telecom-OS-and-Robo
Dataset Description
The architecture addresses a structural gap in modern AI and autonomous systems: the separation between computation and external consequence. Existing protocols and controls (identity, access control, encryption, logging, policy engines) govern movement, authentication, and recording of data. They do not, by themselves, make the transition from a generated act to an externally effective act a protected technical precondition.
This dataset provides a clean… See the full description on the dataset page: https://huggingface.co/datasets/sangamdas/Execution-Finality-Security-for-Agentic-AI-Autonomous-Systems-Cloud-Payments-Telecom-OS-and-Robo.africa-synth-laboratory-information-systems-all
Laboratory Information Systems | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-laboratory-information-systems-all.africa-agriculture-and-agrifood-systems-value
Agriculture and agrifood systems — Value | Africa (FAOSTAT) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-agriculture-and-agrifood-systems-value.albanian-medical-exams-systems-mcq-400
Albanian Medical Exams MCQ Dataset
This dataset contains 400 multiple-choice questions from Albanian medical exams, specifically under Fondet e pyetjeve >> Profili Biologji >> Sistemet sections, available here: https://qsha.gov.al/porvimi-i-informatizuar-i-mjekesise/.
Usage
This dataset can be used for training and evaluating question-answering systems in the medical domain for the Albanian language.
License
The questions in this dataset have been extracted… See the full description on the dataset page: https://huggingface.co/datasets/marjpri/albanian-medical-exams-systems-mcq-400.africa-faostat-employment-indicators-agriculture-and-agrifood-systems-oea
Employment Indicators: Agriculture and agrifood systems — Africa (FAOSTAT) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-faostat-employment-indicators-agriculture-and-agrifood-systems-oea.
