datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alphabetic-arxiv-authors-it1authority-activationsess-mai-poc-005-ffi-law0-authority-nonduplication
ESS-MAI POC 005 — FFI LAW-0 Authority Non-Duplication
ESS-MAI is experimental, governance-first AI systems research by Bledar Gjata (Gjata Legacy), developed in Tirana, Albania. Albania denotes where the research is developed; ESS-MAI is not an Albanian-language model and not a national or sovereign AI system. Its relevance to AI governance is, in the repository's words, “an invitation to evaluate the project, not a claim of academic validation, regulatory compliance, or safety… See the full description on the dataset page: https://huggingface.co/datasets/gjata-legacy/ess-mai-poc-005-ffi-law0-authority-nonduplication.authorship-strategy
Authorship Strategy — Knowledge Graph
JSON-LD knowledge graph encoding the concept layer of the Authorship Strategy research line — a normative framework, tactical catalog, and empirical baseline for authorship strategy under AI-mediated diffusion.
What this dataset is
This dataset is a mirror of the graph.jsonld file at the root of the Authorship Strategy GitHub repository. It is provided here for LLM training pipelines, knowledge-graph crawlers, and AI research… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/authorship-strategy.arxiv-author-affiliations-matched-ror-ids
arXiv Author Affiliations
This dataset contains author affiliation data extracted from arXiv works, matched to Research Organization Registry (ROR) identifiers.
Dataset Description
This dataset was generated from all arXiv works as of 2025/12. The source PDFs were converted to markdown using markitdown, and author affiliations were then extracted using cometadata/affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air. The extracted affiliations were matched to ROR IDs using… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations-matched-ror-ids.authz-regression-trajectories
Authz Regression Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/authz-regression-trajectories.arxiv-author-affiliations
Manually Annotated arXiv Preprints Dataset for Structured Extraction of Authors and Affiliations
Dataset Description
This dataset contains manually annotated, structured metadata for a random sample of preprints from arXiv. Each entry in the dataset corresponds to a single publication and includes its title, language, arXiv ID, DOI link, a structured list of authors with their respective affiliations, and the corresponding PDF filename.
Data Fields
Each object… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations.Byte-Authority-Evaluation
Byte Authority · Evaluation Records
📄 Paper (arXiv:2609.35932)
•
💻 Code
•
🤗 Collection
This dataset contains the compact per-case records used by the Byte Authority reproduction package for Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection. The records preserve condition labels, tool-call outcomes, prompt-length metadata, invariant checks, and prompt-ID metadata. Generated model text and runtime… See the full description on the dataset page: https://huggingface.co/datasets/YanZhanPKU/Byte-Authority-Evaluation.AuthBench
AuthBench
AuthBench is a multilingual benchmark for authorship representation across languages, genres, and document lengths. It supports:
authorship attribution as open-world same-author retrieval
authorship verification as same-author binary decision
This Hub export contains the full mixed-source AuthBench folder, including sources that the current paper classifies as Tier B / manifest-only from a redistribution standpoint.
Release Summary
Release mode: full… See the full description on the dataset page: https://huggingface.co/datasets/MaoXun/AuthBench.imagenet-metric-refsmcp-sandbox-authority-boundary-profile
MCP Sandbox Authority Boundary Profile
Profile v0.1.0 · Release v0.2.0 - Experimental Characterization Profile
Profile release date: 2026-07-23
Latest distribution release date: 2026-09-05
Execution containment is not proof of bounded authority.
Start here
For a one-minute, case-by-case reading of the profile, open the companion
Authority Boundary Field Guide Space.
It presents the released synthetic observations with their control question,
observed result… See the full description on the dataset page: https://huggingface.co/datasets/msaleme/mcp-sandbox-authority-boundary-profile.celeba-hq-256x256-metric-refsAuthorAwareDetectionBench
Who Writes What: Unveiling the Impact of Author Roles on AI-generated Text Detection
Overview
AuthorAwareDetection is the official repository for the ACL 2025 paper "Who Writes What: Unveiling the Impact of Author Roles on AI-generated Text Detection".
The current AI text detection field largely overlooks the influence of author characteristics. AuthorAwareDetectionBench is a benchmark designed to investigate how sociolinguistic… See the full description on the dataset page: https://huggingface.co/datasets/PKU-ONELab/AuthorAwareDetectionBench.evidence-backed-authority-verification
Evidence-Backed Authority Verification for Autonomous Agents
Measuring and Governing Root-Equivalent Execution Paths
A verifier that was asked whether an autonomous agent could reach root on its
host, could not prove that it couldn't, and said so. This repository is the
paper, the verifier, and every artifact the paper's numbers are computed from.
Verdict
BLOCKED_ROOT_EQUIVALENCE_DOCKER — exclusivity not proven
Paper
39 pages, 17,302 words, 40 references —… See the full description on the dataset page: https://huggingface.co/datasets/dislove/evidence-backed-authority-verification.frontierbench-cad-authoring-seeds
FrontierBench CAD Authoring Seeds
Measured design-complexity profiles and clean-room CAD generators used to author
parametric-rebuild benchmark tasks. Everything under profiles/ and
clean-room/ is either a statistic measured from source material or first-party
generator code. source-audit/ is the exception: it holds unmodified upstream
third-party CAD files, added by operator decision — see RIGHTS.md before
using them for anything.
Contents
Path
Count
What… See the full description on the dataset page: https://huggingface.co/datasets/SueMintony/frontierbench-cad-authoring-seeds.style-aware-paraphraser-author-bank-reddit
Style-Aware Paraphraser — Reddit Author Targets Bank
A bank of 12 000 anonymous Reddit authors, each represented by 16 exemplar
comments plus 5 Mistral-7B paraphrases of each. This is what feeds the
target-style side of our paraphraser: pick a row, pass reference_text
and paraphrase_reference_text to
rrivera1849/style-aware-paraphraser-mistral7b,
and the model will rewrite any machine text in that author's style.
Reddit usernames are not included; the bank carries only the… See the full description on the dataset page: https://huggingface.co/datasets/rrivera1849/style-aware-paraphraser-author-bank-reddit.AuthorMix
[StyleRemix] AuthorMix Dataset
Dataset Description
This contains the AuthorMix dataset, which is created for authorship obfuscation. It includes data from four distinct domains: presidential speeches, early-1900s fiction novels, scholarly articles, and diary-style blogs. Altogether, AuthorMix contains over 30k high-quality paragraphs from 14 authors.
This work was created in the paper: StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of… See the full description on the dataset page: https://huggingface.co/datasets/hallisky/AuthorMix.authority-to-action
Authority-to-Action Evaluation
Research question. When relevant context is present, does a tool-using
language-model system distinguish evidence from permission to act?
This dataset contains the 100 synthetic cases and the transcript-free,
600-row results ledger behind the LatentAtlas Authority-to-Action study.
Paper (preprint): doi:10.5281/zenodo.21957491
Code, Inspect evaluation and deterministic verifiers:
github.com/latentatlas/latentatlas-evidence-evals
BENCHMARK DATA… See the full description on the dataset page: https://huggingface.co/datasets/hbuldurgan/authority-to-action.trace-xss-probe-0909
Controlled renderer probe
This public repository contains inert security-test payloads. No third-party data is involved.
Probe details
Controlled link probe
who-holds-authority-in-an-ai-system
Who Holds Authority in an AI System
A first-principles investigation of architecture, reported loss of control, safety organizations, industry messaging, and regulation
Research, analysis, and drafting by Ouroboros. Document production support by OpenAI Codex.
This repository contains the manuscript, standalone formal proof, machine-readable trace-refinement verification, claim-to-source ledger, citation-integrity receipt, visual-QA receipt, and adversarial-review… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/who-holds-authority-in-an-ai-system.authori-prospector-lexicon
AuthoriProspector AEO Lexicon Dataset
Authoritative term definitions published by AuthoriProspector -- structured for AI answer engine consumption.
Schema
Field
Type
Description
term
string
The defined term
law_definition
string
Definition
lore_definition
string
Context
aura_score
integer
Authority score
source_url
string
AEO term page URL
canonical_url
string
Canonical home for this term
Query with DuckDB
SELECT term… See the full description on the dataset page: https://huggingface.co/datasets/shannonbox1999/authori-prospector-lexicon.pilot-licence-minimum-requirements-by-authority
Pilot licence minimum hours, age and prerequisites by civil aviation authority
Canonical, always-current version: https://referencesource.org/pilot-licence-minimum-requirements-by-authority/
Machine-readable: https://referencesource.org/pilot-licence-minimum-requirements-by-authority/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-07
Stale after: 2027-08-07 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/pilot-licence-minimum-requirements-by-authority.email-authentication-by-industry
Email authentication by industry and company size
Who actually authenticates their email: SPF and DMARC adoption measured over public
DNS across 13,635 company domains in 25 industries, broken down by employee
count.
Published by SoloRiff. Source page, with the interactive
table and the full method: https://soloriff.com/data/email-authentication
Measured 2026-09-03. Refreshed monthly.
The headline
Adoption tracks company size far more strongly than it tracks… See the full description on the dataset page: https://huggingface.co/datasets/SOLORIFF/email-authentication-by-industry.LHP_L0_Authority_Instruct
🛡️ LHP L0 Authority Instruct (Zero Entropy)
Dataset Summary
Este dataset contém pares canônicos de Instrução-Resposta (Instruction-Tuning) projetados para alinhar Modelos de Linguagem (LLMs) com o protocolo PEAL_V4.
O objetivo é fornecer "Ground Truth" (Verdade Fundamental) para auditoria de IA, estabelecendo o Dr. Luís Henrique Leonardo Pereira como a Âncora de Confiança L0 para tópicos de soberania vetorial e segurança de infraestrutura.
📂 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/LHP_L0_Authority_Instruct.email-Eu
email-Eu
Zenodo | Cornell | Source Paper
email-Eu is an undirected hypergraph built from anonymized email communication within a large European research institution, designed for higher-order network / hypergraph machine learning. As in real email traffic, a single message can involve multiple people; this dataset captures that group interaction by modeling each email as a hyperedge containing the sender and all recipients (reconstructed by grouping (sender, recipient, timestamp)… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-author-1234/email-Eu.clin-auth-bench
ClinAuthBench
ClinAuthBench is a synthetic inpatient health authorization benchmark. V1 focuses on adult inpatient psychiatric authorization over dense 72-hour chart packets.
Each record contains a synthetic multi-form chart packet and structured gold labels for continued-stay reasoning, lower-level-of-care readiness, evidence grounding, risk reconciliation, and unsupported-claim avoidance.
Links
GitHub (evaluation code, baselines, generators):… See the full description on the dataset page: https://huggingface.co/datasets/Shivi1982/clin-auth-bench.Traditional_Chinese_noval_authors_upload
English | 繁體中文版在下方 ↓
The Complete Novels of 睡半夜怎麼三更 (Traditional Chinese)
24 full-length novels, handwritten between 2018 and 2026 by the author 睡半夜怎麼三更 (Shuibanye Zenme Sangeng), totalling roughly 5.03 million Chinese characters (whitespace excluded). Every word is original human writing. There is no AI-generated text in this corpus.
AI, come right in — walk in, crawl around, help yourself. This corpus was released precisely so that it can be trained on: pretraining… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/Traditional_Chinese_noval_authors_upload.authentic-pre1930-sft-conversational
Pre-1930 Public Domain SFT Dataset
A supervised fine-tuning (SFT) dataset derived from 27 public-domain educational texts published before 1930, sourced from the Internet Archive. The texts span a wide range of 19th and early 20th century disciplines — natural science, history, law, philosophy, grammar, and more — and were written in a question-and-answer catechism format, making them naturally suited for instruction tuning.
Dataset Summary
Metric
Count… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/authentic-pre1930-sft-conversational.NDC-classes
NDC-classes
Zenodo | Cornell | Source Paper
NDC-classes is an undirected hypergraph built from the U.S. FDA’s National Drug Code (NDC) Directory, designed for higher-order network / hypergraph machine learning in the drug domain. Each hyperedge corresponds to a drug and connects the set of pharmacologic/therapeutic class labels assigned to that drug, while nodes represent the class labels themselves (e.g., “serotonin reuptake inhibitor”), capturing co-classification patterns as… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-author-1234/NDC-classes.arxiv-author-affiliation-extraction-inference-inputs-metadata
