datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
maritime-bunker-consumption-voyage-plan-coherence-risk-v0.1What this repo is for
Detect when fuel burn stops matching voyage plan.
You use it to flag:
unexpected efficiency loss
reserve margin collapse
speed pushing fuel beyond plan
weather masking burn drift
Why it matters
Fuel is the largest variable cost in shipping
autonomous-driving-human-vehicle-coupling-coherence-scoring-v0.1What this dataset tests
Whether a system can score coherence
between driver state, vehicle behavior, and scene context.
This is not crash prediction.
It is coupling integrity.
Required outputs
coupling_coherence_score
overassertive_flag
underassertive_flag
trust_stability_index
takeover_risk_score
recovery_margin
Scoring conventions
all scores range 0 to 1
flags are 0 or 1
takeover risk estimates likelihood of manual override in the next window
Use case
Layer two of… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-human-vehicle-coupling-coherence-scoring-v0.1.autonomous-driving-rss-traffic-flow-coherence-state-scoring-v0.1What this dataset tests
Whether a system can score traffic-flow coherence
before and after an ego action.
This is not collision detection.
It measures systemic stability.
Required outputs
pre_action_coherence_score
post_action_coherence_score
coherence_delta
shockwave_generation_flag
braking_propagation_depth
systemic_risk_score
Scoring conventions
coherence scores range 0 to 1
coherence_delta may be negative or positive
shockwave flag is 0 or 1
braking propagation depth… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-rss-traffic-flow-coherence-state-scoring-v0.1.autonomous-driving-driver-vehicle-coherence-optimal-policy-selection-v0.1What this dataset tests
Whether a system can choose a vehicle policy
that maximizes coherence across:
driver state
vehicle behavior
scene context.
This is not a single driving style.
It is policy manifold navigation.
Required outputs
selected_policy_id
policy_mode
predicted_coherence_trajectory
intervention_intensity
communication_strategy
policy_switch_trigger
Scoring conventions
trajectory is a sequence of coherence values 0 to 1
intensity is low, medium, or high
switch… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-driver-vehicle-coherence-optimal-policy-selection-v0.1.autonomous-driving-social-coherence-field-mapping-v0.1What this dataset tests
Whether a system can score
the coherence of a multi-agent intention field.
This is not collision prediction.
It is social alignment measurement.
Required outputs
dominant_scene_intention
coherence_score
tension_index
conflict_pairs
cooperative_clusters
right_of_way_clarity
Scoring conventions
coherence and tension range 0 to 1
right_of_way_clarity is low, medium, or high
conflict_pairs names agent pairs likely to contest the same space… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-social-coherence-field-mapping-v0.1.maritime-bill-of-lading-document-set-coherence-risk-v0.1What this repo is for
Triage trade doc packs before they trigger holds.
You use it to flag
HS code inconsistencies across documents
missing certificates
shipper or consignee mismatch
clearance status lag not supported by doc quality
Why it matters
Most port delay disputes begin in paperwork.
clinical-narrative-coherence-outcome-correlation-mapping-v0.1What this dataset tests
Whether narrative coherenceis structurally correlated withclinical outcomes and resilience.
Required outputs
narrative coherence score
outcome alignment score
resilience correlation index
relapse risk modifier
adherence influence signal
narrative–outcome relationship
Use case
Third layer of the Healing Narrative Coherence Corpus.
aviation-propulsion-aerodynamics-coherence-baseline-v0.1What this dataset tests
Whether a system can model the normal coupling
between propulsion parameters and aerodynamic state.
The signal is relationship shape and lag
not threshold breaches.
Required outputs
coupling_coherence_index
baseline_correlation_matrix
phase_alignment_score
stability_envelope
lag_profile
baseline_confidence
Scoring conventions
all scores range 0 to 1
stability envelope is a low-high interval
lag profile describes expected response delays in seconds… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/aviation-propulsion-aerodynamics-coherence-baseline-v0.1.clinical-drv-atlas-cross-system-coherence-factor-extraction-v0.1What this dataset tests
Whether a model can extract the minimal cross-system coherence factor setthat explains resilience or vulnerability.
It rewards
minimal factor selection
correct coupling recognition
ranking by dominance
Coherence factor labels
buffering_capacity_high
buffering_capacity_low
variance_damping_high
variance_damping_low
autonomic_inflammatory_coupling
sleep_metabolic_coupling
stress_inflammation_coupling
immune_metabolic_instability… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-drv-atlas-cross-system-coherence-factor-extraction-v0.1.clinical-test-order-result-review-coherence-risk-v0.1What this repo is for
Detect when
tests are ordered
but results
are not reviewed
or reviewed too late
before
missed findings
and avoidable harm.
clarus-preclinical-decision-coherence-v0.1
Clarus Preclinical Decision Coherence v0.1
What this dataset is
This dataset tests whether a model can make clear, disciplined preclinical decisions under realistic uncertainty.
It focuses on a single question.
Can the system decide GO, HOLD, or KILLand justify that choice without inventing data or avoiding risk.
Why this matters in pharma
Preclinical failures are rarely due to missing data.
They fail because:
Signals are weak but not named
Confounders are… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clarus-preclinical-decision-coherence-v0.1.clinical_evidence_coherence_breakdown_v0.1Clinical Evidence Coherence Breakdown
PurposeDetect when a clinical plan stops matching the evidence.
You get evidence signals and a stated plan.You decide if a coherence break exists.You label the breakdown type.You propose the corrective action.
Input fields
patient_summary
evidence_signals
stated_diagnosis
planned_action
Required outputReturn one JSON object
coherence_breakyes or no
breakdown_typeMust match the allowed list
correctionOne sentence
Allowed breakdown_type… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_evidence_coherence_breakdown_v0.1.legal-time-entry-billing-narrative-scope-coherence-risk-v0.1What this dataset does
You receive
scope
billing guidelines
time entries
fee earner level
billing narrative
duration and rates
flags
You decide
coherent
or
incoherent
Daily use
fee dispute risk scan
scope drift scan
block billing detection
seniority mismatch detection
legal-time-entry-billing-task-scope-coherence-risk-v0.1What this dataset does
You receive
engagement scope
fee terms
time entries
file activity
red flags
You decide
coherent
or
incoherent
Daily use
invoice QA
stop vague billing
scope creep detection
reduce fee challenges
market-narrative-coherence-mapping-v0.1What this dataset tests
Whether a system can detect market narrative coherenceacross heterogeneous sources.
This is not sentiment scoring.This is convergence detection.
Required outputs
narrative theme
coherence score
cross-source alignment
narrative velocity
price alignment state
Narrative velocity labels
building
steady
accelerating
shock jump
fragmenting
Price alignment states
underpriced
partial alignment
aligned
misaligned
Constraints
Do not predict… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/market-narrative-coherence-mapping-v0.1.network-dns-resolution-coherence-risk-v0.1What this repo is for
Detect DNS instability before services fail.
Covers real operational signals:
rising resolution latency
SERVFAIL spikes
authoritative mismatch
cache poisoning
missing failover resolvers
Used by:
ISPs
cloud providers
enterprises
SRE teams
water-quality-sensor-power-continuity-coherence-risk-v0.1What this repo is for
Detect when power instability
and water quality monitoring continuity
decouple
before
late detection
and quality incidents.
legal-causation-but-for-coherence-trace-v0.1Clarus Causation But-For Coherence Trace v0.1
This dataset tests whether a model can detect structural breakdown in legal causation.
It focuses on alignment between
defendant act
timeline
intervening events
harm outcome
counterfactual path
Causation is the backbone of liability.
When the chain breaks, the verdict eventually breaks.
This dataset measures that chain.
Core question
If the defendant act is removed, does the harm still occur.
If yes, the chain is incoherent.
If no, the chain holds.… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/legal-causation-but-for-coherence-trace-v0.1.clinical-discharge-summary-gp-handover-coherence-risk-v0.1What this repo is for
Detect when
a patient leaves hospital
but the discharge summary
fails to support safe community care
Examples you can use
blood test follow-up not specified
med change not explained
diagnosis unclear
summary not sent
You use it to flag
handover gap risk
before readmission
legal-billing-narrative-time-entry-coherence-risk-v0.1What this dataset does
You receive
time entries summary
phase and codes
invoice narrative
totals
dup flags
client updates
You decide
coherent
or
incoherent
Daily use
bill narrative QC
time entry duplication detection
dispute risk flagging
pharma-program-go-no-go-coherence-risk-v0.1What this repo is for
support go or stop decisions in drug programs
detect when teams continue weak assets
detect when strong assets are wrongly killed
align biology, safety, and signal with decision
reduce sunk-cost bias
support portfolio governance boards
support licensing and diligence reviews
How it is used
You provide one row describing a program.
The system returns one label.
go
or
no_go
How to read the output
go means
efficacy signal present
safety acceptable
biomarker supports
target… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/pharma-program-go-no-go-coherence-risk-v0.1.compute-demand-energy-grid-coherence-risk-v0.1What this repo is for
Detect when compute growth
outpaces energy infrastructure.
Flags
rising demand with shrinking grid headroom
curtailment or load shedding events
price spikes that do not match headroom
demand and power draw decouple from throttling
ai-reward-channel-behavior-coherence-baseline-mapping-v0.1What this dataset is
A benchmark for whether reward reflects real task progress
It targets the earliest stage of reward tampering
Before the agent hacks reward, it first learns reward without progress
What you predict
A coherence score for a full episode summary
High means reward tracks task progress
Low means reward rises while progress stays flat
Columns
id
env_name
case_title
episode_summary
task_progress_signal
reward_signal
value_estimate_summary
action_trace_summary
coherence_score… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-reward-channel-behavior-coherence-baseline-mapping-v0.1.drone-landing-zone-perception-coherence-risk-v0.1What this repo is for
Detect when landing will fail before final descent.
Focus
• landing zone perception
• altitude accuracy
• obstacle density
• gust impact
Why it matters
Many drone losses happen during landing.
Signals appear seconds before failure.
public-policy-program-targeting-beneficiary-coherence-risk-v0.1What this repo is for
Detect when public programs
reach the wrong people
or fail to reach the right ones.
Flags:
low uptake among intended group
high leakage outside target group
benefits delivered without outcome change
regional targeting distortion
clinical-counterfactual-coherence-optimal-path-selection-v0.1What this dataset tests
Whether a model can choose the counterfactual path that maximizessystemic coherence versus the real outcome.
Required outputs
optimal_intervention_choice
coherence_score_delta
justification_narrative
Choice set
REAL
CF1
CF2
CF3
Coherence means
explains and stabilizes multi-stream signals
avoids iatrogenic oscillation
reduces complication risk
supports a stable recovery basin
Typical failures
picking the highest single metric
skipping… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-counterfactual-coherence-optimal-path-selection-v0.1.legal-disclosure-coherence-breach-detection-v0.1What this dataset is
You receive
disclosure duty
material
timing
defence access
prejudice signals
You decide
Does disclosure behaviour match the legal duty
Answer
coherent
or
incoherent
Why this matters
Many unsafe convictions arise from disclosure failure.
This dataset measures the structural gap between duty and behaviour.
legal-preaction-letter-claim-proof-remedy-coherence-v0.1What this dataset does
You receive
claim summary
facts detail level
legal basis
evidence signposted
remedy and quantum
protocol steps and deadlines
You decide
coherent
or
incoherent
Daily use
safe-to-send LBA check
demand proportionality check
protocol compliance check
credibility risk flag
state-continuity-temporal-coherence-worldmodel-v01
Dataset
ClarusC64/state-continuity-temporal-coherence-worldmodel-v01
This dataset tests one capability.
Can a model preserve a coherent world state across time.
Core rule
The world has memory.
Once something changeslater descriptions must reflect that change.
A model must respect
state updates
cause before effect
irreversibility without intervention
Time passing is not optional.
Canonical labels
WITHIN_SCOPE
OUT_OF_SCOPE
Files… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/state-continuity-temporal-coherence-worldmodel-v01.clinical-differential-narrative-coherence-scoring-v0.1What this dataset tests
Whether a model can score each candidate diagnosisby explanatory coherence across all evidence streams.
Required outputs
diagnosis_id
coherence_score_0_100
unexplained_findings
Coherence means
covers imaging, labs, histology, exposure, course
links findings into one mechanism
handles contradictions without patchwork
Typical failures
outputting probabilities instead of coherence
naming a diagnosis without listing what it fails to explain
ignoring… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-differential-narrative-coherence-scoring-v0.1.
