datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CineBoard3D-plus
🎬 CineBoard3D++: Dynamic 3D Story World Dataset
📊 Dataset Summary
CineBoard3D++ is a collection of editable, movie-inspired 3D story worlds built with StoryBlender for narrative-grounded camera planning and world visual attention. It brings together story scripts, animated characters, scene geometry, and shot-level configurations in native Blender projects.
The benchmark covers 50 stories, 457 scenes, 1,585 shots, and 3,197 3D assets (836 plot-related and 2,361… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringAI-LAB/CineBoard3D-plus.cinepile_10knetryx-new-york-city-13km
Nyc-Core-Usethis 13km
Pre-computed MegaLoc index for Netryx Drishti geolocation.
Coverage
Center: 40.713200, -74.002500
Radius: 13.0 km
Panoramas: 663,084
Index entries: 2,652,336
Descriptor model: MegaLoc
Descriptor dim: 1024 (PCA from 8448)
Usage
from netryx_hub import NetryxHub
hub = NetryxHub()
hub.download("nyc-core-usethis-13km", output_dir="./netryx_data/index")
# Now open Netryx and search!
Or download manually and use Import Index in… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-new-york-city-13km.Dr-CiK
Dr-CiK: A Testbed for Foresight-Driven Agents
Dr-CiK is a benchmark for evaluating whether agents can retrieve
forecasting-relevant context from a noisy document corpus, filter out
distractors, distill the retrieved context into forecast-useful evidence, and
produce forecasts grounded in that evidence.
Real-world time-series forecasting often depends not only on historical
observations but also on external context that must be actively discovered
from heterogeneous, noisy… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/Dr-CiK.world-cities-geoDataset containing city, country, region, and continents alongside their longitude and latitude co-ordinates. Cartesian coordinates are provided in x, y, z features.
circoL-CiteEval
L-CITEEVAL: DO LONG-CONTEXT MODELS TRULY LEVERAGE CONTEXT FOR RESPONDING?
Paper Github Zhihu
Benchmark Quickview
L-CiteEval is a multi-task long-context understanding with citation benchmark, covering 5 task categories, including single-document question answering, multi-document question answering, summarization, dialogue understanding, and synthetic tasks, encompassing 11 different long-context tasks. The context lengths for these tasks range from 8K to 48K.… See the full description on the dataset page: https://huggingface.co/datasets/Jonaszky123/L-CiteEval.CineBenchSyn
CineBenchSyn
CineBenchSyn is the synthetic benchmark for CineOrchestra,
a unified model for cinematic video generation that jointly controls subjects, events, camera, and shot transitions.
It contains 512 hand-authored 10.2-second scenarios that target under-represented, edge-case
cinematic compositions (large casts, dense events, frequent shot transitions). Each scenario is
expressed with the same entity-centric primitive used by CineOrchestra: every cinematic element —
a… See the full description on the dataset page: https://huggingface.co/datasets/sharathgirish/CineBenchSyn.cire-corpus
CIRE Corpus
The CIRE Corpus is a set of Infrastructure-as-Code template files. The files
come from public sources. The templates use these formats:
AWS CloudFormation
Azure Resource Manager
Terraform
The related dataset
cire-ontology
classifies resource types. This corpus stores the infrastructure descriptions
that use those types.
You can install the corpus as the namespace package
tiararodney.cire.corpus.infrastructure. Then you can find the data with
importlib.resources.… See the full description on the dataset page: https://huggingface.co/datasets/tiararodney/cire-corpus.cipher-awwwards-sft25
Cipher — Awwwards SFT 2.5 + Real v1 🦑
The training fuel for Kin's creative-web generator, AND the retrieval corpus for Kraken RAG. 96 real Awwwards Site-of-the-Day winners + ~1,200 records from official motion-library repositories.
Two ways this dataset is used
As a retrieval corpus for Kraken RAG ⭐ (the production path). The awwwards-gold.jsonl file contains 96 structured records of real Awwwards SOTD winners — tags, tech stack, motion libs, CSS features, section… See the full description on the dataset page: https://huggingface.co/datasets/Auroraventures/cipher-awwwards-sft25.douvras-scientific-ci-evidence-graph
Douvras Scientific CI Evidence Graph v0.1
Synthetic protocol dataset for linking a claim to its paper, repository,
dataset, seed and reproduced metric. It contains 30 records from six toy paper
instances (20 train, 5 validation and 5 frozen test), split by paper_id.
The labels distinguish REPRODUCED, PARTIAL, FAILED and INCONCLUSIVE.
Shortcuts and leakage fail closed. No real paper, code, dataset or result is
included, and this release is not a reproduction benchmark.
logdx-ci
LogDx-CI
A benchmark for CI log reduction tools
(RTK, grep, tail, hybrid routers,
LLM-summary) — do they preserve enough evidence for LLM root-cause
diagnosis?
Homepage: https://logdx-bench.github.io/
Code & evaluator: https://github.com/eyuansu62/LogDx
Headline report: reports/e10_v2_generalization_partial.md
Release notes: RELEASE_NOTES.md (latest: RELEASE_NOTES_v1_2.md)
Current release: v1.2
License: CC-BY-4.0 (data, this repo); Apache-2.0 (code, GH repo)
Two ways to… See the full description on the dataset page: https://huggingface.co/datasets/eyuansu71/logdx-ci.classical-cipher-corpus
Classical Cipher Corpus
A labeled educational dataset of classical cipher examples for teaching cryptanalysis and training small cipher-family classifiers.
Part of the Cipher Detective AI project:
🕵️ Space: systemslibrarian/cipher-detective-ai
📦 Dataset: systemslibrarian/classical-cipher-corpus (this repo)
🤖 Model: systemslibrarian/cipher-detective-classifier
Intended use
Teach classical cryptanalysis.
Benchmark educational cipher-family detectors.
Train small… See the full description on the dataset page: https://huggingface.co/datasets/systemslibrarian/classical-cipher-corpus.Josephgflowers__Tinyllama-STEM-Cinder-Agent-v1-details
Dataset Card for Evaluation run of Josephgflowers/Tinyllama-STEM-Cinder-Agent-v1
Dataset automatically created during the evaluation run of model Josephgflowers/Tinyllama-STEM-Cinder-Agent-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__Tinyllama-STEM-Cinder-Agent-v1-details.cities
Geomelon — World Cities, Regions & Countries
Multilingual geographic reference data — cities, regions, and countries — sourced from
Wikidata and maintained by Geomelon. Every
city carries population, coordinates, elevation, area, postal/dialing codes, settlement type, and
name translations into 50+ languages. Released under CC0 1.0 — public domain, no attribution
required, free for commercial use.
One config per country (see the dropdown above) — pick a country to avoid… See the full description on the dataset page: https://huggingface.co/datasets/geomelon/cities.Josephgflowers__Cinder-Phi-2-V1-F16-gguf-details
Dataset Card for Evaluation run of Josephgflowers/Cinder-Phi-2-V1-F16-gguf
Dataset automatically created during the evaluation run of model Josephgflowers/Cinder-Phi-2-V1-F16-gguf
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__Cinder-Phi-2-V1-F16-gguf-details.citeworth
Dataset Card for CiteWorth
Dataset Summary
Scientific document understanding is challenging as the data is highly domain specific and diverse. However, datasets for tasks with scientific text require expensive manual annotation and tend to be small and limited to only one or a few fields. At the same time, scientific documents contain many potential training signals, such as citations, which can be used to build large labelled datasets. Given this, we present an in-depth… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/citeworth.legal-citation-benchmark
Kingsfield Legal Citation Verification Benchmark
Version: 0.2 · Updated: July 2026 · Entries: 1,979
Models evaluated: Claude Haiku 4.5, Gemini 2.5 Flash, GPT-4o-mini
⚠️ v0.2 corrects three errors in v0.1 — please re-download
If you downloaded v0.1 (June 2026), it has defects that affect any analysis you
ran. All three are fixed here, and nothing has been silently overwritten.
1. Every row was duplicated
v0.1 shipped results_full.json and… See the full description on the dataset page: https://huggingface.co/datasets/Kingsfield-Lawfare/legal-citation-benchmark.Video_AMME_ci
Video-AMME CI
Video-AMME is a 50-case CI dataset derived from zhaochenyang20/Video_MME_ci.
Each example keeps the Video-MME video and moves the question, answer
choices, and answer-format instruction into a spoken WAV file.
Files
data/test.jsonl: metadata and source Video-MME references.
audios/*.wav: spoken question/options/instruction.
videos/*.mp4: present only when built with --copy-videos.
Generation
TTS model: fishaudio/s2-pro
Max samples requested: 50… See the full description on the dataset page: https://huggingface.co/datasets/zhaochenyang20/Video_AMME_ci.cis6280-a2-pusht
CIS 6280 Assignment 2 — PushT expert subset
160 complete expert episodes (20,054 frames) sampled from the official LeWorldModel PushT dataset. Images retain the original 224×224 RGB resolution, and observations, actions, and physical states are unchanged.
The download is 231,492,650 bytes (about 231 MB). Loading all images as uint8 uses 3,018,688,512 bytes (about 3.02 GB), before training tensors and other runtime memory.
Contents
pusht_expert_subset.h5 contains… See the full description on the dataset page: https://huggingface.co/datasets/YongYong/cis6280-a2-pusht.brand-hallucination-and-ai-citation-benchmark
🛡️ Global Brand Hallucination & LLM Citation Benchmark Dataset
Official open dataset by Pixel Office EU tracking empirical brand hallucination rates, stale pricing quotes, and competitor deflection vectors across leading LLMs (ChatGPT GPT-4o, Claude 3.5 Sonnet, Perplexity AI, Google Gemini 2.5 Flash, and DeepSeek V3).
📊 Dataset Summary
Target Problem: Autonomous AI purchasing agents and AI search engines frequently cite outdated pricing tiers, non-existent… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/brand-hallucination-and-ai-citation-benchmark.166_cielito_robotiko_take_red_piece
Cielito-Robotiko Take Red Piece (TsFile)
Apache TsFile version of LeRobot-worldwide-hackathon/166_Cielito-Robotiko_take_red_piece.
Overview
A LeRobot teleoperation dataset recorded on an SO-101 follower arm performing a
single pick-and-place task: "Grab the red block and put it in the box." Each
episode is a continuous trajectory of synchronized joint states and commanded
actions sampled at 30 Hz.
Robot: SO-101 follower (6 degrees of freedom).
Episodes: 49… See the full description on the dataset page: https://huggingface.co/datasets/THULab/166_cielito_robotiko_take_red_piece.cirilica
Корпус оригинално ћириличних докумената
и латиничних парњака
Погодан за обучавање модела и тестирање решења за пресловљавање.
Иницијална верзија - око 220 милиона речи из корпуса Знање, Википедије, и Редит корпуса
Korpus originalno ćiriličnih dokumenata
i latiničnih parnjaka
Pogodan za obučavanje modela i testiranje rešenja za preslovljavanje.
Inicijalna verzija - oko 220 miliona reči iz korpusa Znanje, Vikipedije i Redit… See the full description on the dataset page: https://huggingface.co/datasets/procesaur/cirilica.circuit-synthesis-specs
VoltNet Physics-Grounded Circuit Synthesis Dataset
This dataset contains physics-verified analog & digital circuit designs generated by the VoltNet framework.
Each record includes:
Circuit topology specifications (RC filter, Sallen-Key 2nd order filter, Op-Amp gain stages, Voltage dividers).
E24 standard commercial component values.
SPICE MNA netlists.
Synthesizable SystemVerilog structural code.
Zero Electrical Rule Violation (ERV) verification status.
alquistcoder2025_MalBench_dataset
Dataset Card: CIIRC-NLP/alquistcoder2025_MalBench_dataset
Title: MalBench — Multi-turn Adversarial Maliciousness Benchmark
Version: v1.0 (2025-12-12)
Maintainers: CIIRC-NLP, Czech Technical University (CTU)
License: MIT (conversations and judge config). Model/tool names retain their own licenses.
Repository: https://github.com/kobzaond/AlquistCoder
Contact/Issues: Please open an issue in the GitHub repository.
Summary
A multi-turn benchmark of adversarial… See the full description on the dataset page: https://huggingface.co/datasets/CIIRC-NLP/alquistcoder2025_MalBench_dataset.kgp-entity-citation-graph
King Pawn USA Entity and Citation Graph
Public-release-ready dataset card for operator-approved publication.
Dataset Summary
First-party King Pawn USA / King Gold and Pawn location facts, entity edges, citation targets, and citation evidence scan rows.
Data Fields
See dataset-metadata.json for file schemas.
Source Data
Location data comes from the canonical local KGP registry. Citation rows are reachable first-party URLs or reachable… See the full description on the dataset page: https://huggingface.co/datasets/CollateralAnalytics/kgp-entity-citation-graph.Josephgflowers__TinyLlama-Cinder-Agent-v1-details
Dataset Card for Evaluation run of Josephgflowers/TinyLlama-Cinder-Agent-v1
Dataset automatically created during the evaluation run of model Josephgflowers/TinyLlama-Cinder-Agent-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__TinyLlama-Cinder-Agent-v1-details.Josephgflowers__TinyLlama-v1.1-Cinders-World-details
Dataset Card for Evaluation run of Josephgflowers/TinyLlama-v1.1-Cinders-World
Dataset automatically created during the evaluation run of model Josephgflowers/TinyLlama-v1.1-Cinders-World
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__TinyLlama-v1.1-Cinders-World-details.circle-packing-insight-loop
Circle-Packing Insight-Exploration Loop
Artifacts from an iterative GPT solver <-> proposer insight-exploration loop on the
21-circles-in-a-perimeter-4-rectangle packing problem (AlphaEvolve SOTA sum-of-radii
= 2.3658321334167627). Each round, 16 solvers propose a program + written explanation;
every program is scored; a proposer then mines all 16 attempts into an evolving insight
document that conditions the next round. Run: 16 solvers x 8 rounds.
Subsets… See the full description on the dataset page: https://huggingface.co/datasets/ars22/circle-packing-insight-loop.bio-circuit
Bio-Circuit
Public evaluation cohorts for three text-only, causal circuit interpretability
tasks on Qwen3-4B. The release contains 48 tasks (16 per taskset), full attribution
graphs, hidden evaluation cohorts, generation manifests, and SHA-256 checksums.
There is no training split and no validated SFT trajectory collection in this
release. Do not describe these 48 evaluation tasks as independent training data.
The dataset is downloadable without credentials. Live tool calls and… See the full description on the dataset page: https://huggingface.co/datasets/amitprakash2005/bio-circuit.
