datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Shamela4_Full_DB
Shamela 4 — Full Islamic Library Corpus
A complete extraction of al-Maktaba al-Shamela (الشاملة) v4, containing 8,589 books across 40 categories of classical Islamic sciences. Extracted from the original Lucene + Sqlite Shamela DB on 2026-04-26 with ~7.6 million pages and ~19 GB of Arabic text.
Dataset Structure
stage0_raw/
├── _meta/ # Cross-cutting metadata (Parquet + JSONL)
│ ├── extraction_manifest.json # Global extraction record
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AuthenticIlm/Shamela4_Full_DB.alphabetic-arxiv-authors-it1single-author-arxiv
Single-author arXiv Computer Science
Metadata for arXiv records classified in Computer Science that list exactly one author.
default retains the original daily-file import. fast stores historical data
in monthly files and adds new submissions as daily update files; it is the
configuration used by the public archive because it makes filtering much faster.
This dataset contains metadata only. arXiv is the source of truth; use each record's
arxiv_url and pdf_url to read the paper.
authority-activationsspooky-author-identificationguardian_authorshipA dataset cross-topic authorship attribution. The dataset is provided by Stamatatos 2013.
1- The cross-topic scenarios are based on Table-4 in Stamatatos 2017 (Ex. cross_topic_1 => row 1:P S U&W ).
2- The cross-genre scenarios are based on Table-5 in the same paper. (Ex. cross_genre_1 => row 1:B P S&U&W).
3- The same-topic/genre scenario is created by grouping all the datasts as follows.
For ex., to use same_topic and split the data 60-40 use:
train_ds = load_dataset('guardian_authorship', name="cross_topic_<<#>>",
split='train[:60%]+validation[:60%]+test[:60%]')
tests_ds = load_dataset('guardian_authorship', name="cross_topic_<<#>>",
split='train[-40%:]+validation[-40%:]+test[-40%:]')
IMPORTANT: train+validation+test[:60%] will generate the wrong splits because the data is imbalanced
* See https://huggingface.co/docs/datasets/splits.html for detailed/more examplesblog_authorship_corpusThe Blog Authorship Corpus consists of the collected posts of 19,320 bloggers gathered from blogger.com in August 2004. The corpus incorporates a total of 681,288 posts and over 140 million words - or approximately 35 posts and 7250 words per person.
Each blog is presented as a separate file, the name of which indicates a blogger id# and the blogger’s self-provided gender, age, industry and astrological sign. (All are labeled for gender and age but for many, industry and/or sign is marked as unknown.)
All bloggers included in the corpus fall into one of three age groups:
- 8240 "10s" blogs (ages 13-17),
- 8086 "20s" blogs (ages 23-27),
- 2994 "30s" blogs (ages 33-47).
For each age group there are an equal number of male and female bloggers.
Each blog in the corpus includes at least 200 occurrences of common English words. All formatting has been stripped with two exceptions. Individual posts within a single blogger are separated by the date of the following post and links within a post are denoted by the label urllink.
The corpus may be freely used for non-commercial research purposes.50k_persian_poem_authorauthority-provenance
authority-provenance
A per-verse authority provenance surface for the Hebrew Bible and New Testament. For every verse it
records independent signals bearing on the authority of the text at that point: textual stability (is the
reading secure in the critical text?), compositional attribution (who wrote it, and on what evidence?),
and canonical reception (how the church received it). These axes are kept separate so that questions of
manuscript evidence, authorship, and reception… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/authority-provenance.ALIA_mixed_authentic_synthetic_MT
Dataset Card for ALIA_mixed_authentic_synthetic_MT
Dataset Summary
Large-scale multilingual parallel corpus covering English and Spanish paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish–Hindi and Spanish–Chinese portions of the dataset include synthetic Spanish translations generated from English using… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA_mixed_authentic_synthetic_MT.arxiv-author-affiliations-matched-ror-ids
arXiv Author Affiliations
This dataset contains author affiliation data extracted from arXiv works, matched to Research Organization Registry (ROR) identifiers.
Dataset Description
This dataset was generated from all arXiv works as of 2025/12. The source PDFs were converted to markdown using markitdown, and author affiliations were then extracted using cometadata/affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air. The extracted affiliations were matched to ROR IDs using… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations-matched-ror-ids.ResearchArcade-openreview-authorsblog_authorship_corpusczech_corpus_authorship_recognition
Czech Authorship Recognition Corpus (Kala)
Popis datasetu
Tento dataset byl vytvořen v rámci diplomové práce zaměřené na automatické rozpoznání autorství českých textů. Obsahuje české publicistické texty získané z veřejně dostupných online zdrojů a připravené pro experimenty v úlohách:
přiřazení autorství (authorship attribution)
ověřování autorství (authorship verification)
shlukování podle autorství (authorship clustering)
Zdrojová data
Do… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/czech_corpus_authorship_recognition.contractbench
ContractBench
A deterministic benchmark for measuring observation-contract compliance in LLM agents: whether agents preserve the temporal validity and byte-level integrity of intermediate tool outputs (presigned URLs, OAuth state parameters, JWT tokens, HMAC-protected webhooks, rate-limit windows, etc.).
This dataset is the companion to the NeurIPS 2026 Evaluations & Datasets Track submission. It contains two complementary artifacts in a single repository:
Subfolder
What's… See the full description on the dataset page: https://huggingface.co/datasets/nips26-anon-author/contractbench.ARES-Bench
ARES-Bench
ARES-Bench is the open audit substrate released with the paper
Auditing LLM User Simulators for Recommender A/B Testing (NeurIPS 2026, ED
Track, under review). It turns the ARES reliability-audit view — the LLM
backbone is the measurement instrument under test, not an interchangeable
implementation detail — into a reproducible protocol over structured
behavioral logs, a portable visual sandbox, and a screenshot cache.
This release hosts the 17,000-session core corpus that… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-ares-authors/ARES-Bench.ResearchArcade-openreview-papers-authorsauthorship-verification
Dataset Card for Dataset Name
Dataset for authorship verification, comprised of 12 cleaned, modified, open source authorship verification and attribution datasets.
Dataset Details
Code for cleaning and modifying datasets can be found in https://github.com/swan-07/authorship-verification/blob/main/Authorship_Verification_Datasets.ipynb and is detailed in paper.
Datasets used to produce the final dataset are:
Reuters50
@misc{misc_reuter_50_50_217,
author = {Liu… See the full description on the dataset page: https://huggingface.co/datasets/swan07/authorship-verification.iclr-reproducibility-bundle
Reproducibility Bundle — ICLR 2027 Submission
Evaluation datasets, scripts, and the model checkpoint for the anonymous
ICLR 2027 submission.
Bundle Layout
Directory
Contents
expression_datasets/
Three .h5ad files (GEO-OmicsQA, Human Diseases, Tabula Sapiens)
embeddings/
Pre-computed OmicsLM input vectors (.npz)
geo_omics_qa/
GEO-OmicsQA question files (six .json variants, 3,000 questions)
omicslm_checkpoint/
OmicsLM model checkpoint in HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/Authors2027/iclr-reproducibility-bundle.ess-mai-poc-005-ffi-law0-authority-nonduplication
ESS-MAI POC 005 — FFI LAW-0 Authority Non-Duplication
ESS-MAI is experimental, governance-first AI systems research by Bledar Gjata (Gjata Legacy), developed in Tirana, Albania. Albania denotes where the research is developed; ESS-MAI is not an Albanian-language model and not a national or sovereign AI system. Its relevance to AI governance is, in the repository's words, “an invitation to evaluate the project, not a claim of academic validation, regulatory compliance, or safety… See the full description on the dataset page: https://huggingface.co/datasets/gjata-legacy/ess-mai-poc-005-ffi-law0-authority-nonduplication.Explore-Execute-Chain-Datasetsimagenetauthor_profilinghe corpus for the author profiling analysis contains texts in Russian-language which labeled for 5 tasks:
1) gender -- 13530 texts with the labels, who wrote this: text female or male;
2) age -- 13530 texts with the labels, how old the person who wrote the text. This is a number from 12 to 80. In addition, for the classification task we added 5 age groups: 1-19; 20-29; 30-39; 40-49; 50+;
3) age imitation -- 7574 texts, where crowdsource authors is asked to write three texts:
a) in their natural manner,
b) imitating the style of someone younger,
c) imitating the style of someone older;
4) gender imitation -- 5956 texts, where the crowdsource authors is asked to write texts: in their origin gender and pretending to be the opposite gender;
5) style imitation -- 5956 texts, where crowdsource authors is asked to write a text on behalf of another person of your own gender, with a distortion of the authors usual style.Swedish_Work_environment_Authority
[!NOTE]
Dataset origin: https://portulanclarin.net/repository/browse/parallel-texts-from-swedish-work-environment-authority-processed/7404236aa58b11eaae0e02420a000403bd13d9138a904f33980bd63233eb90bc/
Description
This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu.
Parallel texts from the Swedish… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Swedish_Work_environment_Authority.image-authenticity-battle
Image Authenticity Battle Dataset
This dataset contains real and synthetic/tampered images for human perception studies on AI-generated content detection.
Dataset Structure
Total Images: 19500
Categories: Real, Synthetic (Fully AI-generated), Tampered (AI-edited)
Models: Nano Banana, Qwen, Flux, SD3
Metadata Fields
Each image has the following metadata:
filename: Path to image file
dataset: Source dataset name
category: real/synthetic/tampered… See the full description on the dataset page: https://huggingface.co/datasets/TheGarlic/image-authenticity-battle.authorship-strategy
Authorship Strategy — Knowledge Graph
JSON-LD knowledge graph encoding the concept layer of the Authorship Strategy research line — a normative framework, tactical catalog, and empirical baseline for authorship strategy under AI-mediated diffusion.
What this dataset is
This dataset is a mirror of the graph.jsonld file at the root of the Authorship Strategy GitHub repository. It is provided here for LLM training pipelines, knowledge-graph crawlers, and AI research… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/authorship-strategy.arxiv-author-affiliations
Manually Annotated arXiv Preprints Dataset for Structured Extraction of Authors and Affiliations
Dataset Description
This dataset contains manually annotated, structured metadata for a random sample of preprints from arXiv. Each entry in the dataset corresponds to a single publication and includes its title, language, arXiv ID, DOI link, a structured list of authors with their respective affiliations, and the corresponding PDF filename.
Data Fields
Each object… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations.celeba-hq-256x256authz-regression-trajectories
Authz Regression Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/authz-regression-trajectories.tabgen
