Team Ai
22 results

auth

AuthenticIlm /Shamela4_Full_DB Shamela 4 — Full Islamic Library Corpus A complete extraction of al-Maktaba al-Shamela (الشاملة) v4, containing 8,589 books across 40 categories of classical Islamic sciences. Extracted from the original Lucene + Sqlite Shamela DB on 2026-04-26 with ~7.6 million pages and ~19 GB of Arabic text. Dataset Structure stage0_raw/ ├── _meta/ # Cross-cutting metadata (Parquet + JSONL) │ ├── extraction_manifest.json # Global extraction record │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AuthenticIlm/Shamela4_Full_DB.text-generation10M<n<100M30 likes40k downloads5mo agoHugging Facekalomaze /alphabetic-arxiv-authors-it1text100K<n<1M0 likes7.7k downloads1y agoHugging Facemainakmanna /single-author-arxiv Single-author arXiv Computer Science Metadata for arXiv records classified in Computer Science that list exactly one author. default retains the original daily-file import. fast stores historical data in monthly files and adds new submissions as daily update files; it is the configuration used by the public archive because it makes filtering much faster. This dataset contains metadata only. arXiv is the source of truth; use each record's arxiv_url and pdf_url to read the paper. text100K<n<1M0 likes2.8k downloads2mo agoHugging Facelasrprobegen /authority-activationstext100K<n<1M0 likes2.5k downloads11mo agoHugging Facehkadxqq /spooky-author-identificationtext10K<n<100K0 likes1.1k downloads4y agoHugging FaceEfstathios /guardian_authorshipA dataset cross-topic authorship attribution. The dataset is provided by Stamatatos 2013. 1- The cross-topic scenarios are based on Table-4 in Stamatatos 2017 (Ex. cross_topic_1 => row 1:P S U&W ). 2- The cross-genre scenarios are based on Table-5 in the same paper. (Ex. cross_genre_1 => row 1:B P S&U&W). 3- The same-topic/genre scenario is created by grouping all the datasts as follows. For ex., to use same_topic and split the data 60-40 use: train_ds = load_dataset('guardian_authorship', name="cross_topic_<<#>>", split='train[:60%]+validation[:60%]+test[:60%]') tests_ds = load_dataset('guardian_authorship', name="cross_topic_<<#>>", split='train[-40%:]+validation[-40%:]+test[-40%:]') IMPORTANT: train+validation+test[:60%] will generate the wrong splits because the data is imbalanced * See https://huggingface.co/docs/datasets/splits.html for detailed/more examplestext-classification1K<n<10K6 likes818 downloads3y agoHugging Face