Team Ai
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agentlans /en-document-classification English Document Classification Dataset This dataset provides a curated subset of the first 1 million rows from the allenai/c4 (English configuration), enriched with multi-perspective topic annotations. It is designed for researchers exploring document classification, domain adaptation, and label noise in massive web-crawled corpora. Dataset Summary The dataset integrates predictions from classification models to provide a holistic view of each document’s content.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-classification.texttext-classification1M<n<10M1 likes301 downloads1mo agoHugging Face02asahi417 /multi-domain-document-classification multi_domain_document_classification Multi-domain document classification datasets. Biomedical: chemprot, rct-sample Computer Science: citation_intent, sciie Customer Review: amcd, yelp_review Social Media: tweet_eval_irony, tweet_eval_hate, tweet_eval_emotion The yelp_review dataset is randomly downsampled to 2000/2000/8000 for test/validation/train. chemprot citation_intent hyperpartisan_news rct_sample sciie amcd yelp_review tweet_eval_irony tweet_eval_hate… See the full description on the dataset page: https://huggingface.co/datasets/asahi417/multi-domain-document-classification.text10K<n<100K0 likes237 downloads4y agoHugging Face03agentlans /en-document-topic-classification English Document Topic Classification Dataset English-language web pages classified by document topic, designed to train robust text classifiers and provide ready-to-use data for specific web topics. Purpose: Train generalized document classifiers or extract clean, single-topic corpora for specific downstream tasks. Configurations: Each document topic is available in its own dedicated dataset configuration (e.g., HomeGardening, GamesRecreation). Splits: The All configuration… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-topic-classification.text1M<n<10M0 likes182 downloads1mo agoHugging Face04agentlans /en-document-format-classification English Document Format Classification Dataset English-language web pages classified by document type, designed to train robust text classifiers and provide ready-to-use data for specific web formats. Purpose: Train generalized document classifiers or extract clean, single-format corpora for specific downstream tasks. Configurations: Each document type is available in its own dedicated dataset configuration (e.g., TutorialHow-ToGuide, PersonalAboutPage). Splits: The All… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-document-format-classification.text100K<n<1M0 likes162 downloads1mo agoHugging Face05nutrientdocs /document-classification-benchmark Document Classification Benchmark (open-vocab, zero-shot) Given a document image and an arbitrary set of text labels, which one is right? A held-out, zero-shot, open-vocabulary evaluation for document-type classification — labels are supplied at inference, not baked into a head. Test split only; not for training. Every image is drawn from a permissively-licensed, redistributable source. Powers the document-classification-leaderboard and evaluates document-classification-v2… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/document-classification-benchmark.imagezero-shot-image-classification1K<n<10K0 likes93 downloads2mo agoHugging Face06hf-tuner /rvl-cdip-document-classification rvl-cdip-document-classification This dataset is created from original aharley/rvl_cdip dataset using this notebook Dataset Summary This dataset consists of 8992 grayscale images in 16 classes, with 562 images per class. There are 8000 training images(500 image per class) and 992 test images(62 images per class). The images are sized so their largest dimension does not exceed 1000 pixels. imageimage-classification1K<n<10K0 likes76 downloads1y agoHugging Face07agentlans /multilingual-document-classification Multilingual Document Classification Dataset This dataset contains 100,000 text passages across 100 non-English language-script pairs sourced from the agentlans/HuggingFaceFW-finetranslations-100-languages-sample collection. Each original text passage is paired with its English translation and has been programmatically annotated with domain, writing genre, and educational classifications to facilitate cross-lingual classification and domain adaptation tasks. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/multilingual-document-classification.texttext-classification100K<n<1M0 likes44 downloads5mo agoHugging Face08arbml /Document_Classificationtext100K<n<1M0 likes22 downloads4y agoHugging Face09supergoose /flan_combined_task420_persent_document_sentiment_classificationtext10K<n<100K0 likes17 downloads2y agoHugging Face10Zanzibara1961 /document_classification_reu_ml_school_hwtext10K<n<100K0 likes14 downloads2y agoHugging Face11yjmsvma /document_classification_dataset_v1text1K<n<10K0 likes11 downloads2y agoHugging Face12NorskHelsenett /LOS_Document_classification_ETI LOS Document Classification (ETI) Norwegian public-sector documents labelled with their top-level LOS category (Felles vokabular / common vocabulary, level 1). Built to train a small, self-hostable text classifier by distilling LLM-generated labels. Splits split rows labels use train ~1982 silver — LLM-generated (Claude Haiku 4.5, few-shot) training test 551 human gold — manually annotated honest evaluation The train labels are weak supervision… See the full description on the dataset page: https://huggingface.co/datasets/NorskHelsenett/LOS_Document_classification_ETI.texttext-classification1K<n<10K0 likes11 downloads4mo agoHugging Face13N1CKNGUYEN /document_classification_level1_finetune Dataset Card for "document_classification_level1_fintune" More Information needed textn<1K0 likes9 downloads1y agoHugging Face14Hadisawara /indonesian-tax-document-classification Indonesian Tax Document Classification Dataset Dataset Description Dataset ini berisi koleksi sintetis dokumen pajak Indonesia yang digunakan untuk klasifikasi jenis dokumen pajak. Dataset dirancang untuk mendukung penelitian NLP berbahasa Indonesia di bidang administrasi pajak dan pemerintahan daerah. Dataset ini dibuat berdasarkan pengalaman dan pengetahuan dari sistem administrasi pajak daerah (Bapenda), dengan struktur yang mencerminkan dokumen-dokumen nyata… See the full description on the dataset page: https://huggingface.co/datasets/Hadisawara/indonesian-tax-document-classification.tabulartext-classification10K<n<100K0 likes7 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.