datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Medical-Data-Referencecassia-reference-data
CASSIA Reference Data
Raw paper supplementary tables and PDFs used to build CASSIA reference-agent
knowledge docs and to score held-out validation papers.
This is the archive of source materials — the processed reference docs
themselves live in the CASSIA Python package
(CASSIA_python/CASSIA/agents/reference_agent/references/).
Layout
Each top-level directory is one benchmark combo (paper pair = dev set + held-out
validation set). Inside each combo:
<combo>/
├──… See the full description on the dataset page: https://huggingface.co/datasets/elliotxie/cassia-reference-data.recachedpo-reference-data
ReCacheDPO reference data
Generated reference images and cached text-conditioning tensors for ReCacheDPO scheduling experiments. No backbone model weights are included.
The taoquan-2.1 bank contains SD3.5 references; sd3_shift3_train_512_20260918 contains SD3 shift=3 training references. Each bank includes its original generation_config.json and prompts.txt with the exact generation settings and prompt order.
Files are partitioned into part-XX directories. To restore a bank… See the full description on the dataset page: https://huggingface.co/datasets/cachedf/recachedpo-reference-data.pyOpenFOAM-reference-data
pyOpenFOAM Reference Data & Validation Results
OpenFOAM-13 reference simulation data and pyOpenFOAM validation results for pyOpenFOAM — a pure Python/PyTorch reimplementation of OpenFOAM with GPU acceleration and automatic differentiation.
Dataset Summary / 数据摘要
Property
Value
Total reference cases
257
Validated cases
233 (90.7%)
Categories
21
Source
OpenFOAM v11/v13
Reference data size
2.42 GB
pyOpenFOAM results
3.3 MB
Field files analyzed… See the full description on the dataset page: https://huggingface.co/datasets/AlanZee/pyOpenFOAM-reference-data.HUB_reference_dataset
Introduction
This dataset is used for HUB, please consider the cahnges in the pull request. Alternatively use the fork I have crated HUB-forked
pyOpenFOAM-reference-data
pyOpenFOAM Reference Data & Validation Results
OpenFOAM-13 reference simulation data and pyOpenFOAM validation results for pyOpenFOAM — a pure Python/PyTorch reimplementation of OpenFOAM with GPU acceleration and automatic differentiation.
Dataset Summary / 数据摘要
Property
Value
Total reference cases
257
Validated cases
233 (90.7%)
Categories
21
Source
OpenFOAM v11/v13
Reference data size
2.42 GB
pyOpenFOAM results
3.3 MB
Field files analyzed… See the full description on the dataset page: https://huggingface.co/datasets/isunme/pyOpenFOAM-reference-data.Cell_SEQR_Reference_Datasetsscancestry-reference-data
scAncestry Reference Panel
Reference data for scancestry, a tool for inferring genetic ancestry from single-cell genomics data.
Reference genome build: GRCh38 / hg38.
Contents
This dataset bundles imputation, phasing, and population-reference files used by the scAncestry pipeline:
gnomad.genomes.v3.1.2.hgdp_tgp.miss0.01.maf0.01.vcf.gz (+ .tbi) — gnomAD v3.1.2 HGDP+1000G common variants, used for PCA reference… See the full description on the dataset page: https://huggingface.co/datasets/powellgenomicslab/scancestry-reference-data.music-tempo-and-key-reference-data
BPM Interval and Camelot Reference Data
This public dataset contains two reference assets maintained by j022315051:
bpm-interval-reference.csv: BPM values with milliseconds per beat and the calculation used.
camelot.js: major and minor pitch-class mappings to Camelot codes.
These reference assets support the browser-based music tools published at Tap BPM Now.
Scope and boundaries
The release contains reference tables and mapping code only. It does not contain… See the full description on the dataset page: https://huggingface.co/datasets/j022315051/music-tempo-and-key-reference-data.paymind-reference-data
PayMind Reference Dataset
Synthetic/reference payment-routing data for PayMind, an open-source payment route intelligence engine.
This dataset is designed to demonstrate PayMind's training, evaluation, and routing workflow across route selection, transaction reliability, and expected settlement time.
Important: This dataset contains synthetic/reference data only. It does not contain real customers, real transactions, payment credentials, personally identifiable information, or… See the full description on the dataset page: https://huggingface.co/datasets/navk8690/paymind-reference-data.Anime_Character_Transfer_and_Reference_Dataset
Anime Character Transfer and Reference Dataset
This dataset is designed for anime-style character transfer, reference-based image editing, and multi-reference character consistency experiments.
Each sample contains a source/reference pair and a text prompt. Most samples also include a generated target image and a metadata file. A small number of samples are kept as reference-only entries, so they can still be used for reference-pair tasks or future target completion.… See the full description on the dataset page: https://huggingface.co/datasets/LAXMAYDAY/Anime_Character_Transfer_and_Reference_Dataset.ai-data-governance-reference-cases
ADGL Reference Cases and Governance Profiles
This Dataset repository accompanies the AI Data Governance Layer (ADGL) public research project.
ADGL models Knowledge Governance → Analysis Governance → Consequence Governance, with INFORM, DECIDE, and ACT as principal consequence dispositions and Audit + Provenance spanning the complete governance trajectory.
Configurations
reference_cases: eight structured reference cases with policy, input fixture, and expected… See the full description on the dataset page: https://huggingface.co/datasets/GBSNResearch/ai-data-governance-reference-cases.genomicsops-reference-datadeveloper-reference-datasets
Developer Reference Datasets
Open, reproducible lookup tables that web and app developers reach for constantly — computed from first principles, not scraped, so every value is exact and re-runnable. CC BY 4.0.
Quick answers (straight from the data)
What is 16:9 in pixels? 1920×1080, 1280×720, 3840×2160. 9:16 (Stories, Reels, TikTok) is those flipped. → aspect-ratios, resolutions
What contrast ratio does WCAG require? 4.5:1 for normal text (AA), 3:1 for large… See the full description on the dataset page: https://huggingface.co/datasets/cleanorlabs/developer-reference-datasets.reference-free-rl-summarization-data
Reference-free RL Summarization Experimental Data
This repository contains experimental data splits, metadata, and processed subsets used for a study on verifier-composable penalty-shaped reinforcement learning for reference-free summarization.
Configs
vnexpress: Vietnamese VnExpress train/validation/test split used in the study. Unless explicit redistribution permission is available, this config releases metadata and split information only.
cnn_dailymail_subset:… See the full description on the dataset page: https://huggingface.co/datasets/phuongntc/reference-free-rl-summarization-data.reference-datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 3,
"total_frames": 597,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rhecker/reference-dataset.reference-sft-dataset
Reference SFT Dataset
A curated, deduplicated, domain-labeled reference dataset for supervised fine-tuning (SFT) of language models. It merges instruction-following conversations, raw knowledge text, and specialized domain QA into a single 13-column schema with a fixed domain vocabulary and a per-row veracity score.
Rows: 207,424
Domains: 10 (fixed vocabulary)
Sources: 27 upstream datasets
Format: sharded Parquet (2 shards, train-00000-of-00002.parquet /… See the full description on the dataset page: https://huggingface.co/datasets/Minnimaro/reference-sft-dataset.Anime_Character_Transfer_and_Reference_Dataset_3600plus
Anime Character Transfer and Reference Dataset — IDs after 003600
This repository is a locally validated, volume-packaged subset of
LAXMAYDAY/Anime_Character_Transfer_and_Reference_Dataset.
Selection: numeric id > 003600
Samples: 6,056
Source ID range: 003611–011228
Samples with targets: 5,535
Reference-only samples: 521
Tar volumes: 4
Every sample remains complete within one volume. Tar member paths preserve the
source repository's train/ref1, train/ref2, train/prompt… See the full description on the dataset page: https://huggingface.co/datasets/LAXMAYDAY/Anime_Character_Transfer_and_Reference_Dataset_3600plus.bash-reference-manual-general-QAs
Dataset generated from bash reference manual.
book information like date and bash version are available within the very first rows of the dataset
this dataset is pretty small in general, but covering almost all of the definition and technical terms, commands and flags in the book
columns : "Question", "Answer"
emotion_dataset_for_tts_with_transcriptions_and_reference_voice_v1paymind-reference-data-v2
PayMind Synthetic Payment Dataset — V4
Synthetic payment-routing data for developing, training and benchmarking PayMind.
This dataset provides the V4 synthetic training environment for PayMind, an open-source payment intelligence connector.
It is designed for three predictive responsibilities:
Engine
Objective
Candidate Generator
Learn which payment routes fit a transaction
Reliability Engine
Estimate transaction success probability
Settlement Intelligence… See the full description on the dataset page: https://huggingface.co/datasets/navk8690/paymind-reference-data-v2.filtering-reference-databrand-structured-data-reference
Brand Structured Data Reference v1.0
This reference maps common public brand facts to structured data concepts that can help people, search engines, and AI systems understand a brand more clearly.
It is intended for independent brands, small businesses, founder-led companies, service providers, local businesses, and early-stage products that need a clearer public identity online.
This is not a ranking guide and it does not guarantee search visibility, rich results, AI… See the full description on the dataset page: https://huggingface.co/datasets/farosio/brand-structured-data-reference.nomenclature-de-reference-de-la-donnee-services-et-equipements-publics
Nomenclature de référence de la donnée services et équipements publics
[!NOTE]
Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Nomenclature de référence de la donnée services et équipements publics qui est disponible à l'adresse https://www.data.gouv.fr/datasets/648095a9ce7de8cc2409271a
Description
Structuration de la donnée en nomenclature
**Thématique **Services et équipements publics
**Producteur **Brest… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/nomenclature-de-reference-de-la-donnee-services-et-equipements-publics.SPV_MIA_reference_data_tldr_mixtral_8x7bReferenceDataReference_Extraction_Datasacs-a-proces-photos-de-l-affaire-reference-2-b-10795
Sacs à procès - Photos de l'affaire référence 2 B 10795
[!NOTE]
Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Sacs à procès - Photos de l'affaire référence 2 B 10795 qui est disponible à l'adresse https://www.data.gouv.fr/datasets/654c5ad3417f2a304a12b363
Description
Ce jeu de données propose les photos de certaines pièces de l'affaire référencée 2 B 10795 du jeu de données Sacs à procès du Parlement de Toulouse, Lot… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/sacs-a-proces-photos-de-l-affaire-reference-2-b-10795.mining_domain_knowledge_reference_datasetdonnees-changement-climatique-sqr-series-quotidiennes-de-reference
Données changement climatique - SQR (Séries Quotidiennes de Référence)
[!NOTE]
Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Données changement climatique - SQR (Séries Quotidiennes de Référence) qui est disponible à l'adresse https://www.data.gouv.fr/datasets/6569b00fe24fc9e1e482f74e
Description
Présentation
Les Séries Quotidiennes de Référence (SQR) sont une sélection de données climatologiques… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/donnees-changement-climatique-sqr-series-quotidiennes-de-reference.
