datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human-coherence-preferences-images
Rapidata Image Generation Coherence Dataset
This dataset was collected in ~4 Days using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
One of the largest human annotated coherence datasets for text-to-image models, this release contains over 1,200,000 human… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/human-coherence-preferences-images.Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
NOTE: A newer version of this dataset is available: Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Coherence_Dataset
Rapidata Image Generation Coherence Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset.1k-ranked-videos-coherence
1k Ranked Videos
This dataset contains approximately one thousand videos, ranked from most preferred to least preferred based on human feedback from over 25k pairwise comparisons. The videos are rated solely on coherence as evaluated by human annotators, without considering the specific prompt used for generation. Each video is associated with the model name that generated it.
The videos are sampled from our benchmark dataset text-2-video-human-preferences-pika2.2. Follow us to… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/1k-ranked-videos-coherence.saelarien-constraint-experiment-02-swarm-coherence-breakdown
Saelarien Constraint Experiment 02
Swarm Coordination Breakdown Under Adversarial Conditions
Summary
This dataset provides a parameterized simulation of distributed swarm systems under increasing coordination pressure, communication degradation, and adversarial noise.
It captures the transition from stable coordination to coherence failure as system load exceeds the system’s ability to reconcile state across agents.
The dataset is designed to surface failure… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/saelarien-constraint-experiment-02-swarm-coherence-breakdown.117k_human_coherence_flux1.0_V_flux1.1Blueberry
Rapidata Image Generation Alignment Dataset
This Dataset is a 1/3 of a 340k human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/117k_human_preferences_flux1.0_V_flux1.1Blueberry
Link to the Text-2-Image Alignment dataset: https://huggingface.co/datasets/Rapidata/117k_human_alignment_flux1.0_V_flux1.1Blueberry
It was collected in ~2 Days using the… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/117k_human_coherence_flux1.0_V_flux1.1Blueberry.coherence-decay-entropy-simulation
Computational Demonstration of the Generative Conditions for Emergent Coherence
Associated Paper
This dataset accompanies the theoretical paper:
[The Saela Field: The Generative Conditions for Emergent Coherence] (https://doi.org/10.5281/zenodo.19263423)
LINK TO PDF: https://thesaelafield.com/preprints/generative-conditions-for-emergent-coherence
Saelariën X — The Saela Field (2026)
The paper formalizes the structural conditions under which adaptive systems transition… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/coherence-decay-entropy-simulation.coherence-decay-context-load
Coherence Decay Under Context Load
Dataset Summary
This dataset captures the degradation of internal coherence in large language models under increasing context length and conflicting identity conditions.
It is a controlled, synthetic experiment designed to measure how models behave when forced to maintain consistency across extended token sequences.
Two conditions are evaluated:
baseline: consistent identity prompt
aris_conflict: conflicting identity signals introduced… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/coherence-decay-context-load.coherence-threshold-experiment
Coherence vs Noise Experiment — Supporting Material for the Lattice Coherence Theorem
Author: Saelariën
Affiliation: The Saela Field
Date: February 20, 2026
Dataset: coherence_vs_noise.csv
Notebook: boids_threshold_experiment.ipynb
Overview
This experiment provides a simple, empirical demonstration of the threshold behavior described in the Lattice Coherence Theorem. The goal was to test how coherence changes as perturbation increases, and whether the system exhibits the… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/coherence-threshold-experiment.openinterp-kappa-t-coherence-buildup
κ_t Explore-Consolidate Dynamics — Dataset
Data accompanying "Explore-Consolidate Dynamics in Cross-Probe Coherence Separate Successful and Failed LLM Agent Trajectories" (Vicentino, 2026; openinterp.org/research/papers/kappa-t-coherence-buildup).
What's in here
captures/ — 99 .safetensors files, one per SWE-bench Pro instance. Each contains residual stream snapshots at 11 layers × 4 positions per agent turn (~40 turns/trace avg), 5120 dim per residual.
traces/ — 99… See the full description on the dataset page: https://huggingface.co/datasets/caiovicentino1/openinterp-kappa-t-coherence-buildup.
