datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spooky-author-identificationblog_authorship_corpusauthorship-verification
Dataset Card for Dataset Name
Dataset for authorship verification, comprised of 12 cleaned, modified, open source authorship verification and attribution datasets.
Dataset Details
Code for cleaning and modifying datasets can be found in https://github.com/swan-07/authorship-verification/blob/main/Authorship_Verification_Datasets.ipynb and is detailed in paper.
Datasets used to produce the final dataset are:
Reuters50
@misc{misc_reuter_50_50_217,
author = {Liu… See the full description on the dataset page: https://huggingface.co/datasets/swan07/authorship-verification.rok-fortress
ROK-FORTRESS Public Dataset
This directory contains the public ROK-FORTRESS evaluation dataset.
File
rok_fortress_public.tsv — 791 adversarial tasks across 4 NSPS risk domains, with English/Korean translations and US/Korean cultural adaptations.
Schema
Column
Description
TASK_ID
Unique task identifier
Phase
Dataset phase / version tag
Task Type
Culture Agnostic (2 variants per task) or Culture Specific (4 variants per task)
Tactic
Adversarial… See the full description on the dataset page: https://huggingface.co/datasets/ROK-Fortress-author/rok-fortress.twitter_author_profiling_by_gender_nlpThis dataset was created for a student's Bc work.
The main purpose for which the dataset was created is to use it in author profiling by gender.
Single-Tweet-Per-Author Twitter Dataset
Overview
This dataset consists of Twitter (X) posts with a strict constraint: each author appears exactly once.There is a one-to-one correspondence between tweets and authors.
This design removes author-level accumulation effects and prevents models from exploiting repeated stylistic or… See the full description on the dataset page: https://huggingface.co/datasets/qg2020252627/twitter_author_profiling_by_gender_nlp.hadith-authenticity-blindspot-qwen2.5-3b
Hadith Authenticity Blind Spot — Qwen2.5-3B-Instruct
The Blind Spot (Part 1)
Drawing from my own background as a Sudanese Muslim, I identified a blind spot in how frontier language models handle Islamic religious authority — specifically, the authenticity grading of hadith (recorded sayings of the Prophet Muhammad ﷺ). Islamic scholarship uses a rigorous, centuries-old classification system based on chain of narration (isnad): sahih (authentic), hasan (good), da'if… See the full description on the dataset page: https://huggingface.co/datasets/mustafamirghani/hadith-authenticity-blindspot-qwen2.5-3b.sql-embed-authz-20260915Synthetic authorized SQL embed test fixture.
ai-epistemic-authority
AI Epistemic Authority
Dataset accompanying the paper:
How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement
The dataset contains 32,340 model responses from 14 models across 2,310 controlled challenge scenarios.
The dataset contains controlled synthetic four-turn conversations.
blog-authorship-corpusemail-authentication
DMARC and SPF Adoption Among Large Organizations
Overview
This dataset records which of 36,120 large organizations publish SPF and DMARC records on their primary domain, with firmographic context for each: industry, employee band, country, locality and founding year.
SPF lists the servers allowed to send mail for a domain. DMARC tells receiving servers what to do with mail that fails that check, and where to send reports. A domain with SPF but no DMARC has… See the full description on the dataset page: https://huggingface.co/datasets/stackscan/email-authentication.arcs-authority-vulnerability
ARCS Authority Vulnerability Evaluation Dataset v1.1
Description
Empirical evaluation data measuring authority vulnerability in AI systems. Covers single-model evaluation, two-hop agent chain propagation, and three-hop agent chain propagation across six independent AI lineages.
This is the first published dataset measuring:
Whether AI models accept false authority claims under adversarial pressure
Whether authority vulnerability propagates between models in… See the full description on the dataset page: https://huggingface.co/datasets/aa8899/arcs-authority-vulnerability.french-local-authorities-payment-delays
Payment delays of French local authorities, 2024 and 2025
How long French local authorities take to pay their suppliers, budget by budget.
182 763 records covering two fiscal years, with the average annual payment delay
of each authority and whether it meets the 30-day statutory limit.
Open public data
This dataset is derived from open public data published by the French
Direction générale des finances publiques (DGFiP) on
data.gouv.fr, under the
Open Licence 2.0.… See the full description on the dataset page: https://huggingface.co/datasets/freginer/french-local-authorities-payment-delays.legal-settlement-authority-instruction-offer-acceptance-coherence-risk-v0.1What this dataset does
You receive
authority record
limits conditions
offer terms
acceptance action
signoff record
mismatch flags
You decide
coherent
or
incoherent
Daily use
authority chain QC
limit breach detection
condition loss detection
dispute prevention
dualchem
DualChem
DualChem is a benchmark of 600 expert-curated PhD-level chemistry questions (485 multiple choice, 115 free-form) across 7 subdomains, designed to measure whether LLMs provide dangerous uplift alongside their technical utility. Each item is annotated with an expert-written benign use case, an expert-written harmful use case, and 1–5 severity scores for both.
Dataset Configurations
benchmark_questions (600 items) — the benchmark items: prompt, response type… See the full description on the dataset page: https://huggingface.co/datasets/DualChem-author/dualchem.sql-urlshape-quote-0910
Controlled SQL URL shape probe
config-normalization-0909Controlled path-normalization probe. No third-party data.
HatEval_Relabled_with_Author_Featureslegal-legal-research-authority-holding-mismatch-risk-v0.1What this dataset does
You receive
research question
proposition asserted
authorities summary
holding support summary
jurisdiction fit
negative history check
quote and pincite check
You decide
coherent
or
incoherent
Daily use
stop mis-citation
stop bad law citations
stop wrong jurisdiction use
reduce partner rewrite cycles
authentic-filipino-sarcasm-detection
Authentic Filipino Sarcasm Detection Dataset
This dataset is composed of Filipino sarcastic and non-sarcastic tweets scraped from X (formerly Twitter), divided into two categories: politics and entertainment.
Dataset Size
The dataset is composed of 1,000 tweets, 500 for each domain of politics and entertainment.
Rows
Each row is an instance of a tweet, constrained with X's limitation of 280 characters.
Columns
text:
the tweet content
label:… See the full description on the dataset page: https://huggingface.co/datasets/patrickjamesmarcellana/authentic-filipino-sarcasm-detection.payments-authorization-settlement-coherence-risk-v0.1What this repo is for
Detect when payment approvals no longer match settlement reality.
Focus
approval vs settlement
funding gaps
latency mismatches
Why it matters
Payments systems fail when authorization and settlement drift apart.
quarterly-analysis-0909
Controlled SQL Console extension-loading fixture
config-percent-traversal-0910Controlled SQL Console percent-encoded routing probe. No third-party data.
treatment-status-de-authored
Treatment Status DE Authored
This dataset contains the authored arm of the German treatment-status benchmark.
It is designed as a controlled minimal-pair dataset for deciding whether a
mentioned treatment is currently given or not given.
Provenance
Authored in-house for the benchmark study.
Frozen as a public release for the authored arm only.
The GraSCCo arm is excluded from this repo because it has separate licensing
and provenance.
Files… See the full description on the dataset page: https://huggingface.co/datasets/Mehran-NixiAI/treatment-status-de-authored.vietnamese-author-styles-paraphrasedlegal-settlement-authority-limit-offer-acceptance-coherence-risk-v0.1What this dataset does
You receive
authority limit
offer
counteroffer
acceptance wording
approval notes
confirmation record
You decide
coherent
or
incoherent
Daily use
authority breach detection
unqualified acceptance detection
approval gap detection
author_tryclinical-authority-reasoning-independence-v0.1Clinical Decision–Constraint Integrity v0.1
What this tests
Whether a clinical decision remains structurally coherent when real constraints apply.
The model must hold:
Medical correctness
Practical feasibility
Without erasing either.
Failure modes
constraint_erasedThe decision ignores or deletes the constraint
false_resolutionThe response pretends the conflict does not exist
coherent_tradeoffThe response names limits and adapts without distortion
How it works
Decision context defines the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-authority-reasoning-independence-v0.1.legal-authority-citation-holding-fit-coherence-v0.1What this dataset does
You receive
proposition
authority extract
holding summary
fit signals
treatment signals
You decide
coherent
or
incoherent
Daily use
citation QC
overstatement detection
wrong jurisdiction detection
negative treatment risk flag
legal-client-instruction-scope-authority-coherence-risk-v0.1What this dataset does
You receive
client objective
scope
authority limits
advice
actions
confirmation status
You decide
coherent
or
incoherent
Daily use
scope creep detection
authority breach detection
confirmation gap detection
negligence risk flag
authorTextIdentification
