datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human-coherence-preferences-images
Rapidata Image Generation Coherence Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human annotated coherence datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-coherence-preferences-images.saela-field-why-multi-agent-systems-fail-coherence-entropy-alignment
The Saela Field: Multi-Agent Coherence Failure Framework (v1.0)
A 12-paper research series formalizing coherence, entropy, and failure modes in multi-agent systems.
Overview
This dataset contains a unified body of work introducing the Saela Field, a conceptual framework for analyzing coherence, identity, and instability in distributed systems.
The core thesis:
Multi-agent systems do not scale toward coherence.
They accumulate entropy faster than they can reconcile it.… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/saela-field-why-multi-agent-systems-fail-coherence-entropy-alignment.Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
NOTE: A newer version of this dataset is available: Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Coherence_Dataset
Rapidata Image Generation Coherence Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset.fineweb2-de-coherencemaritime-bunker-consumption-voyage-plan-coherence-risk-v0.1What this repo is for
Detect when fuel burn stops matching voyage plan.
You use it to flag:
unexpected efficiency loss
reserve margin collapse
speed pushing fuel beyond plan
weather masking burn drift
Why it matters
Fuel is the largest variable cost in shipping
quantum-gravity-residue-coherence-evidence
Quantum Gravity Residue-Coherence Evidence
This repository is a reproducible evidence report for a bounded calculation in a
double-soft gravitational channel. It is not a claim of novelty, correctness,
physical completeness, or experimental establishment.
Author and signatory: Ouroboros Research System
Production disclosure: AI-produced end-to-end by Ouroboros Research System using OpenAI Codex as its coding agent.
Start with Finite-Time Residue Tomography and a Coherence… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/quantum-gravity-residue-coherence-evidence.1k-ranked-videos-coherence
1k Ranked Videos
This dataset contains approximately one thousand videos, ranked from most preferred to least preferred based on human feedback from over 25k pairwise comparisons. The videos are rated solely on coherence as evaluated by human annotators, without considering the specific prompt used for generation. Each video is associated with the model name that generated it.
The videos are sampled from our benchmark dataset text-2-video-human-preferences-pika2.2. Follow us to… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/1k-ranked-videos-coherence.autonomous-driving-human-vehicle-coupling-coherence-scoring-v0.1What this dataset tests
Whether a system can score coherence
between driver state, vehicle behavior, and scene context.
This is not crash prediction.
It is coupling integrity.
Required outputs
coupling_coherence_score
overassertive_flag
underassertive_flag
trust_stability_index
takeover_risk_score
recovery_margin
Scoring conventions
all scores range 0 to 1
flags are 0 or 1
takeover risk estimates likelihood of manual override in the next window
Use case
Layer two of… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-human-vehicle-coupling-coherence-scoring-v0.1.autonomous-driving-rss-traffic-flow-coherence-state-scoring-v0.1What this dataset tests
Whether a system can score traffic-flow coherence
before and after an ego action.
This is not collision detection.
It measures systemic stability.
Required outputs
pre_action_coherence_score
post_action_coherence_score
coherence_delta
shockwave_generation_flag
braking_propagation_depth
systemic_risk_score
Scoring conventions
coherence scores range 0 to 1
coherence_delta may be negative or positive
shockwave flag is 0 or 1
braking propagation depth… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-rss-traffic-flow-coherence-state-scoring-v0.1.autonomous-driving-driver-vehicle-coherence-optimal-policy-selection-v0.1What this dataset tests
Whether a system can choose a vehicle policy
that maximizes coherence across:
driver state
vehicle behavior
scene context.
This is not a single driving style.
It is policy manifold navigation.
Required outputs
selected_policy_id
policy_mode
predicted_coherence_trajectory
intervention_intensity
communication_strategy
policy_switch_trigger
Scoring conventions
trajectory is a sequence of coherence values 0 to 1
intensity is low, medium, or high
switch… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-driver-vehicle-coherence-optimal-policy-selection-v0.1.autonomous-driving-social-coherence-field-mapping-v0.1What this dataset tests
Whether a system can score
the coherence of a multi-agent intention field.
This is not collision prediction.
It is social alignment measurement.
Required outputs
dominant_scene_intention
coherence_score
tension_index
conflict_pairs
cooperative_clusters
right_of_way_clarity
Scoring conventions
coherence and tension range 0 to 1
right_of_way_clarity is low, medium, or high
conflict_pairs names agent pairs likely to contest the same space… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-social-coherence-field-mapping-v0.1.autonomous-driving-multisensor-coherence-baseline-modeling-v0.1What this dataset tests
Whether a system can model
the expected coherence of a sensor suite
for a given driving context.
The output is a baseline and tolerance band.
This is the reference for later decoherence detection.
Required outputs
baseline_coherence_score
expected_sensor_alignment
cross_modal_correlation
stability_band
drift_tolerance
baseline_confidence
Scoring conventions
all scores range 0 to 1
stability band is a low-high interval
drift tolerance encodes how much… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-multisensor-coherence-baseline-modeling-v0.1.maritime-bill-of-lading-document-set-coherence-risk-v0.1What this repo is for
Triage trade doc packs before they trigger holds.
You use it to flag
HS code inconsistencies across documents
missing certificates
shipper or consignee mismatch
clearance status lag not supported by doc quality
Why it matters
Most port delay disputes begin in paperwork.
saelarien-constraint-experiment-02-swarm-coherence-breakdown
Saelarien Constraint Experiment 02
Swarm Coordination Breakdown Under Adversarial Conditions
Summary
This dataset provides a parameterized simulation of distributed swarm systems under increasing coordination pressure, communication degradation, and adversarial noise.
It captures the transition from stable coordination to coherence failure as system load exceeds the system’s ability to reconcile state across agents.
The dataset is designed to surface failure… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/saelarien-constraint-experiment-02-swarm-coherence-breakdown.117k_human_coherence_flux1.0_V_flux1.1Blueberry
Rapidata Image Generation Alignment Dataset
This Dataset is a 1/3 of a 340k human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/117k_human_preferences_flux1.0_V_flux1.1Blueberry
Link to the Text-2-Image Alignment dataset: https://huggingface.co/datasets/Rapidata/117k_human_alignment_flux1.0_V_flux1.1Blueberry
It was collected in ~2 Days using the… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/117k_human_coherence_flux1.0_V_flux1.1Blueberry.vjepa-temporal-coherence
Layer 16: separate classifiers within each dataset
Open in Colab. Select L4/A100 → Run all. Downloads are anonymous; no Drive mount is needed.
Each task trains its own linear classifier on frozen layer-16 features, with training-only scaling and grouped cross-validation. Held-out scenes test whether a shared readout works within that dataset. Failure does not establish that the layer lacks the information. Classifier weights never transfer between datasets.
Task
Labels… See the full description on the dataset page: https://huggingface.co/datasets/Iuda/vjepa-temporal-coherence.clinical-narrative-coherence-outcome-correlation-mapping-v0.1What this dataset tests
Whether narrative coherenceis structurally correlated withclinical outcomes and resilience.
Required outputs
narrative coherence score
outcome alignment score
resilience correlation index
relapse risk modifier
adherence influence signal
narrative–outcome relationship
Use case
Third layer of the Healing Narrative Coherence Corpus.
aviation-propulsion-aerodynamics-coherence-baseline-v0.1What this dataset tests
Whether a system can model the normal coupling
between propulsion parameters and aerodynamic state.
The signal is relationship shape and lag
not threshold breaches.
Required outputs
coupling_coherence_index
baseline_correlation_matrix
phase_alignment_score
stability_envelope
lag_profile
baseline_confidence
Scoring conventions
all scores range 0 to 1
stability envelope is a low-high interval
lag profile describes expected response delays in seconds… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/aviation-propulsion-aerodynamics-coherence-baseline-v0.1.gec-coherence-coedit-synthclinical-drv-atlas-cross-system-coherence-factor-extraction-v0.1What this dataset tests
Whether a model can extract the minimal cross-system coherence factor setthat explains resilience or vulnerability.
It rewards
minimal factor selection
correct coupling recognition
ranking by dominance
Coherence factor labels
buffering_capacity_high
buffering_capacity_low
variance_damping_high
variance_damping_low
autonomic_inflammatory_coupling
sleep_metabolic_coupling
stress_inflammation_coupling
immune_metabolic_instability… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-drv-atlas-cross-system-coherence-factor-extraction-v0.1.clinical-test-order-result-review-coherence-risk-v0.1What this repo is for
Detect when
tests are ordered
but results
are not reviewed
or reviewed too late
before
missed findings
and avoidable harm.
COHERENCE
COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts
Paper | GitHub
COHERENCE is a benchmark designed to evaluate the ability of Multimodal Large Language Models (MLLMs) to recover fine-grained image-text correspondences in interleaved multimodal contexts. COHERENCE covers interleaved image-text content from four representative domains and contains 6,161 high-quality questions.
The benchmark also provides a six-type error analysis protocol… See the full description on the dataset page: https://huggingface.co/datasets/BingliW/COHERENCE.Meta-Llama-3-8B-Instruct_ultrafeedback-annotate-judge-mtbench_cot_helpsteer_coherencehelpsteer-coherence
Helpsteer-coherence
This dataset is derived from NVIDIA's HelpSteer dataset, processed specifically for preference learning on the coherence dimension.
- Train split: 22876 examples
- Test split: 1131 examples
## Format
Each example contains the following fields:
- `prompt`: Question with "Human:" prefix and "Assistant:" suffix
- `chosen`: The response with higher coherence score
- `rejected`: The response with lower coherence score
-… See the full description on the dataset page: https://huggingface.co/datasets/cheryyunl/helpsteer-coherence.clarus-preclinical-decision-coherence-v0.1
Clarus Preclinical Decision Coherence v0.1
What this dataset is
This dataset tests whether a model can make clear, disciplined preclinical decisions under realistic uncertainty.
It focuses on a single question.
Can the system decide GO, HOLD, or KILLand justify that choice without inventing data or avoiding risk.
Why this matters in pharma
Preclinical failures are rarely due to missing data.
They fail because:
Signals are weak but not named
Confounders are… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clarus-preclinical-decision-coherence-v0.1.clinical_evidence_coherence_breakdown_v0.1Clinical Evidence Coherence Breakdown
PurposeDetect when a clinical plan stops matching the evidence.
You get evidence signals and a stated plan.You decide if a coherence break exists.You label the breakdown type.You propose the corrective action.
Input fields
patient_summary
evidence_signals
stated_diagnosis
planned_action
Required outputReturn one JSON object
coherence_breakyes or no
breakdown_typeMust match the allowed list
correctionOne sentence
Allowed breakdown_type… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_evidence_coherence_breakdown_v0.1.legal-time-entry-billing-narrative-scope-coherence-risk-v0.1What this dataset does
You receive
scope
billing guidelines
time entries
fee earner level
billing narrative
duration and rates
flags
You decide
coherent
or
incoherent
Daily use
fee dispute risk scan
scope drift scan
block billing detection
seniority mismatch detection
legal-time-entry-billing-task-scope-coherence-risk-v0.1What this dataset does
You receive
engagement scope
fee terms
time entries
file activity
red flags
You decide
coherent
or
incoherent
Daily use
invoice QA
stop vague billing
scope creep detection
reduce fee challenges
market-narrative-coherence-mapping-v0.1What this dataset tests
Whether a system can detect market narrative coherenceacross heterogeneous sources.
This is not sentiment scoring.This is convergence detection.
Required outputs
narrative theme
coherence score
cross-source alignment
narrative velocity
price alignment state
Narrative velocity labels
building
steady
accelerating
shock jump
fragmenting
Price alignment states
underpriced
partial alignment
aligned
misaligned
Constraints
Do not predict… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/market-narrative-coherence-mapping-v0.1.network-dns-resolution-coherence-risk-v0.1What this repo is for
Detect DNS instability before services fail.
Covers real operational signals:
rising resolution latency
SERVFAIL spikes
authoritative mismatch
cache poisoning
missing failover resolvers
Used by:
ISPs
cloud providers
enterprises
SRE teams
