Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AuthenticIlm /Shamela4_Full_DB Shamela 4 — Full Islamic Library Corpus A complete extraction of al-Maktaba al-Shamela (الشاملة) v4, containing 8,589 books across 40 categories of classical Islamic sciences. Extracted from the original Lucene + Sqlite Shamela DB on 2026-04-26 with ~7.6 million pages and ~19 GB of Arabic text. Dataset Structure stage0_raw/ ├── _meta/ # Cross-cutting metadata (Parquet + JSONL) │ ├── extraction_manifest.json # Global extraction record │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AuthenticIlm/Shamela4_Full_DB.text-generation10M<n<100M31 likes40k downloads5mo agoHugging Face02kalomaze /alphabetic-arxiv-authors-it1text100K<n<1M0 likes9.1k downloads1y agoHugging Face03mainakmanna /single-author-arxiv Single-author arXiv Computer Science Metadata for arXiv records classified in Computer Science that list exactly one author. default retains the original daily-file import. fast stores historical data in monthly files and adds new submissions as daily update files; it is the configuration used by the public archive because it makes filtering much faster. This dataset contains metadata only. arXiv is the source of truth; use each record's arxiv_url and pdf_url to read the paper. text100K<n<1M0 likes2.4k downloads2mo agoHugging Face04lasrprobegen /authority-activationstext100K<n<1M0 likes2.2k downloads11mo agoHugging Face05hkadxqq /spooky-author-identificationtext10K<n<100K0 likes974 downloads4y agoHugging Face06Efstathios /guardian_authorshipA dataset cross-topic authorship attribution. The dataset is provided by Stamatatos 2013. 1- The cross-topic scenarios are based on Table-4 in Stamatatos 2017 (Ex. cross_topic_1 => row 1:P S U&W ). 2- The cross-genre scenarios are based on Table-5 in the same paper. (Ex. cross_genre_1 => row 1:B P S&U&W). 3- The same-topic/genre scenario is created by grouping all the datasts as follows. For ex., to use same_topic and split the data 60-40 use: train_ds = load_dataset('guardian_authorship', name="cross_topic_<<#>>", split='train[:60%]+validation[:60%]+test[:60%]') tests_ds = load_dataset('guardian_authorship', name="cross_topic_<<#>>", split='train[-40%:]+validation[-40%:]+test[-40%:]') IMPORTANT: train+validation+test[:60%] will generate the wrong splits because the data is imbalanced * See https://huggingface.co/docs/datasets/splits.html for detailed/more examplestext-classification1K<n<10K6 likes878 downloads3y agoHugging Face07barilan /blog_authorship_corpusThe Blog Authorship Corpus consists of the collected posts of 19,320 bloggers gathered from blogger.com in August 2004. The corpus incorporates a total of 681,288 posts and over 140 million words - or approximately 35 posts and 7250 words per person. Each blog is presented as a separate file, the name of which indicates a blogger id# and the blogger’s self-provided gender, age, industry and astrological sign. (All are labeled for gender and age but for many, industry and/or sign is marked as unknown.) All bloggers included in the corpus fall into one of three age groups: - 8240 "10s" blogs (ages 13-17), - 8086 "20s" blogs (ages 23-27), - 2994 "30s" blogs (ages 33-47). For each age group there are an equal number of male and female bloggers. Each blog in the corpus includes at least 200 occurrences of common English words. All formatting has been stripped with two exceptions. Individual posts within a single blogger are separated by the date of the following post and links within a post are denoted by the label urllink. The corpus may be freely used for non-commercial research purposes.text-classification10K<n<100K18 likes734 downloads3y agoHugging Face08Erfan3940 /50k_persian_poem_authortext10K<n<100K0 likes476 downloads10mo agoHugging Face09NuBerea /authority-provenancegated authority-provenance A per-verse authority provenance surface for the Hebrew Bible and New Testament. For every verse it records independent signals bearing on the authority of the text at that point: textual stability (is the reading secure in the critical text?), compositional attribution (who wrote it, and on what evidence?), and canonical reception (how the church received it). These axes are kept separate so that questions of manuscript evidence, authorship, and reception… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/authority-provenance.textfeature-extraction10K<n<100K0 likes424 downloads1mo agoHugging Face10BSC-LT /ALIA_mixed_authentic_synthetic_MT Dataset Card for ALIA_mixed_authentic_synthetic_MT Dataset Summary Large-scale multilingual parallel corpus covering English and Spanish paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish–Hindi and Spanish–Chinese portions of the dataset include synthetic Spanish translations generated from English using… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA_mixed_authentic_synthetic_MT.texttranslation100M<n<1B1 likes387 downloads10mo agoHugging Face11cometadata /arxiv-author-affiliations-matched-ror-ids arXiv Author Affiliations This dataset contains author affiliation data extracted from arXiv works, matched to Research Organization Registry (ROR) identifiers. Dataset Description This dataset was generated from all arXiv works as of 2025/12. The source PDFs were converted to markdown using markitdown, and author affiliations were then extracted using cometadata/affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air. The extracted affiliations were matched to ROR IDs using… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations-matched-ror-ids.texttext-classification1M<n<10M1 likes378 downloads9mo agoHugging Face12ulab-ai /ResearchArcade-openreview-authorstext100K<n<1M0 likes339 downloads7mo agoHugging Face13tasksource /blog_authorship_corpustabular100K<n<1M2 likes335 downloads2y agoHugging Face14MU-NLPC /czech_corpus_authorship_recognition Czech Authorship Recognition Corpus (Kala) Popis datasetu Tento dataset byl vytvořen v rámci diplomové práce zaměřené na automatické rozpoznání autorství českých textů. Obsahuje české publicistické texty získané z veřejně dostupných online zdrojů a připravené pro experimenty v úlohách: přiřazení autorství (authorship attribution) ověřování autorství (authorship verification) shlukování podle autorství (authorship clustering) Zdrojová data Do… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/czech_corpus_authorship_recognition.texttext-classification0 likes286 downloads4mo agoHugging Face15nips26-anon-author /contractbench ContractBench A deterministic benchmark for measuring observation-contract compliance in LLM agents: whether agents preserve the temporal validity and byte-level integrity of intermediate tool outputs (presigned URLs, OAuth state parameters, JWT tokens, HMAC-protected webhooks, rate-limit windows, etc.). This dataset is the companion to the NeurIPS 2026 Evaluations & Datasets Track submission. It contains two complementary artifacts in a single repository: Subfolder What's… See the full description on the dataset page: https://huggingface.co/datasets/nips26-anon-author/contractbench.other1K<n<10K0 likes268 downloads5mo agoHugging Face16neurips2026-ares-authors /ARES-Bench ARES-Bench ARES-Bench is the open audit substrate released with the paper Auditing LLM User Simulators for Recommender A/B Testing (NeurIPS 2026, ED Track, under review). It turns the ARES reliability-audit view — the LLM backbone is the measurement instrument under test, not an interchangeable implementation detail — into a reproducible protocol over structured behavioral logs, a portable visual sandbox, and a screenshot cache. This release hosts the 17,000-session core corpus that… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-ares-authors/ARES-Bench.imageother10K<n<100K2 likes267 downloads5mo agoHugging Face17ulab-ai /ResearchArcade-openreview-papers-authorstext100K<n<1M0 likes264 downloads7mo agoHugging Face18swan07 /authorship-verification Dataset Card for Dataset Name Dataset for authorship verification, comprised of 12 cleaned, modified, open source authorship verification and attribution datasets. Dataset Details Code for cleaning and modifying datasets can be found in https://github.com/swan-07/authorship-verification/blob/main/Authorship_Verification_Datasets.ipynb and is detailed in paper. Datasets used to produce the final dataset are: Reuters50 @misc{misc_reuter_50_50_217, author = {Liu… See the full description on the dataset page: https://huggingface.co/datasets/swan07/authorship-verification.texttext-classification100K<n<1M3 likes247 downloads2y agoHugging Face19Authors2027 /iclr-reproducibility-bundle Reproducibility Bundle — ICLR 2027 Submission Evaluation datasets, scripts, and the model checkpoint for the anonymous ICLR 2027 submission. Bundle Layout Directory Contents expression_datasets/ Three .h5ad files (GEO-OmicsQA, Human Diseases, Tabula Sapiens) embeddings/ Pre-computed OmicsLM input vectors (.npz) geo_omics_qa/ GEO-OmicsQA question files (six .json variants, 3,000 questions) omicslm_checkpoint/ OmicsLM model checkpoint in HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/Authors2027/iclr-reproducibility-bundle.0 likes235 downloads12d agoHugging Face20gjata-legacy /ess-mai-poc-005-ffi-law0-authority-nonduplication ESS-MAI POC 005 — FFI LAW-0 Authority Non-Duplication ESS-MAI is experimental, governance-first AI systems research by Bledar Gjata (Gjata Legacy), developed in Tirana, Albania. Albania denotes where the research is developed; ESS-MAI is not an Albanian-language model and not a national or sovereign AI system. Its relevance to AI governance is, in the repository's words, “an invitation to evaluate the project, not a claim of academic validation, regulatory compliance, or safety… See the full description on the dataset page: https://huggingface.co/datasets/gjata-legacy/ess-mai-poc-005-ffi-law0-authority-nonduplication.textn<1K0 likes227 downloads2d agoHugging Face21anomyous-author /Explore-Execute-Chain-Datasetstext10K<n<100K0 likes225 downloads1y agoHugging Face22gfi-authors /imagenet1M<n<10M0 likes222 downloads21d agoHugging Face23sagteam /author_profilinghe corpus for the author profiling analysis contains texts in Russian-language which labeled for 5 tasks: 1) gender -- 13530 texts with the labels, who wrote this: text female or male; 2) age -- 13530 texts with the labels, how old the person who wrote the text. This is a number from 12 to 80. In addition, for the classification task we added 5 age groups: 1-19; 20-29; 30-39; 40-49; 50+; 3) age imitation -- 7574 texts, where crowdsource authors is asked to write three texts: a) in their natural manner, b) imitating the style of someone younger, c) imitating the style of someone older; 4) gender imitation -- 5956 texts, where the crowdsource authors is asked to write texts: in their origin gender and pretending to be the opposite gender; 5) style imitation -- 5956 texts, where crowdsource authors is asked to write a text on behalf of another person of your own gender, with a distortion of the authors usual style.text-classification10K<n<100K1 likes207 downloads4y agoHugging Face24FrancophonIA /Swedish_Work_environment_Authority [!NOTE] Dataset origin: https://portulanclarin.net/repository/browse/parallel-texts-from-swedish-work-environment-authority-processed/7404236aa58b11eaae0e02420a000403bd13d9138a904f33980bd63233eb90bc/ Description This dataset has been created within the framework of the European Language Resource Coordination (ELRC) Connecting Europe Facility - Automated Translation (CEF.AT) action. For further information on the project: http://lr-coordination.eu. Parallel texts from the Swedish… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Swedish_Work_environment_Authority.0 likes203 downloads2y agoHugging Face25TheGarlic /image-authenticity-battle Image Authenticity Battle Dataset This dataset contains real and synthetic/tampered images for human perception studies on AI-generated content detection. Dataset Structure Total Images: 19500 Categories: Real, Synthetic (Fully AI-generated), Tampered (AI-edited) Models: Nano Banana, Qwen, Flux, SD3 Metadata Fields Each image has the following metadata: filename: Path to image file dataset: Source dataset name category: real/synthetic/tampered… See the full description on the dataset page: https://huggingface.co/datasets/TheGarlic/image-authenticity-battle.imageimage-classification10K<n<100K0 likes199 downloads11mo agoHugging Face26shimo4228 /authorship-strategy Authorship Strategy — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Authorship Strategy research line — a normative framework, tactical catalog, and empirical baseline for authorship strategy under AI-mediated diffusion. What this dataset is This dataset is a mirror of the graph.jsonld file at the root of the Authorship Strategy GitHub repository. It is provided here for LLM training pipelines, knowledge-graph crawlers, and AI research… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/authorship-strategy.tabularn<1K1 likes190 downloads1mo agoHugging Face27cometadata /arxiv-author-affiliations Manually Annotated arXiv Preprints Dataset for Structured Extraction of Authors and Affiliations Dataset Description This dataset contains manually annotated, structured metadata for a random sample of preprints from arXiv. Each entry in the dataset corresponds to a single publication and includes its title, language, arXiv ID, DOI link, a structured list of authors with their respective affiliations, and the corresponding PDF filename. Data Fields Each object… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations.textfeature-extraction1K<n<10K2 likes177 downloads1y agoHugging Face28gfi-authors /celeba-hq-256x25610K<n<100K0 likes176 downloads21d agoHugging Face29rmems /authz-regression-trajectories Authz Regression Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/authz-regression-trajectories.text1K<n<10K0 likes172 downloads16d agoHugging Face30anon-author-tabgen /tabgen0 likes169 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.