datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
en-document-classification
English Document Classification Dataset
This dataset provides a curated subset of the first 1 million rows from the allenai/c4 (English configuration), enriched with multi-perspective topic annotations. It is designed for researchers exploring document classification, domain adaptation, and label noise in massive web-crawled corpora.
Dataset Summary
The dataset integrates predictions from classification models to provide a holistic view of each document’s content.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-classification.multi-domain-document-classification
multi_domain_document_classification
Multi-domain document classification datasets.
Biomedical: chemprot, rct-sample
Computer Science: citation_intent, sciie
Customer Review: amcd, yelp_review
Social Media: tweet_eval_irony, tweet_eval_hate, tweet_eval_emotion
The yelp_review dataset is randomly downsampled to 2000/2000/8000 for test/validation/train.
chemprot
citation_intent
hyperpartisan_news
rct_sample
sciie
amcd
yelp_review
tweet_eval_irony
tweet_eval_hate… See the full description on the dataset page: https://huggingface.co/datasets/asahi417/multi-domain-document-classification.en-document-topic-classification
English Document Topic Classification Dataset
English-language web pages classified by document topic, designed to train robust text classifiers and provide ready-to-use data for specific web topics.
Purpose: Train generalized document classifiers or extract clean, single-topic corpora for specific downstream tasks.
Configurations: Each document topic is available in its own dedicated dataset configuration (e.g., HomeGardening, GamesRecreation).
Splits: The All configuration… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-topic-classification.en-document-format-classification
English Document Format Classification Dataset
English-language web pages classified by document type, designed to train robust text classifiers and provide ready-to-use data for specific web formats.
Purpose: Train generalized document classifiers or extract clean, single-format corpora for specific downstream tasks.
Configurations: Each document type is available in its own dedicated dataset configuration (e.g., TutorialHow-ToGuide, PersonalAboutPage).
Splits: The All… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-format-classification.multilingual-document-classification
Multilingual Document Classification Dataset
This dataset contains 100,000 text passages across 100 non-English language-script pairs sourced from the agentlans/HuggingFaceFW-finetranslations-100-languages-sample collection.
Each original text passage is paired with its English translation and has been programmatically annotated with domain, writing genre, and educational classifications to facilitate cross-lingual classification and domain adaptation tasks.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/multilingual-document-classification.
