Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ranWang /UN_Historical_PDF_Article_Text_Corpus python dataset = load_dataset("ranWang/UN_Historical_PDF_Article_Text_Corpus", split="train") or dataset = load_dataset("ranWang/UN_Historical_PDF_Article_Text_Corpus", split="randomTest") lang_list = ["ar", "en", "es", "fr", "ru", "zh"] for row in dataset: # 获取pdf文章内容 for lang in lang_list: # type == str lang_match_file_content = row[lang] # 如果按页分割 lang_match_file_pages_content = lang_match_file_content.split("\n----\n") text100K<n<1M2 likes1.6k downloads3y agoHugging Face02Karavet /ILUR-news-text-classification-corpus News Texts Dataset We release a dataset of over 12000 news articles from iLur.am, categorized into 7 classes: sport, politics, weather, economy, accidents, art, society. The articles are split into train (2242k tokens) and test sets (425k tokens). For more details, refer to the paper. texttext-classification100K<n<1M8 likes1k downloads4y agoHugging Face03thoshith /hindi-english-raw-text-corpus-uncleanedtext100M<n<1B0 likes985 downloads2y agoHugging Face04zomi-language-corpora /raw-text-corpus 📝 Zomi Raw Text Corpus (Community-Contributed) The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks. This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately. 📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.texttext-generationn<1K0 likes735 downloads5mo agoHugging Face05PersianML /persian-text-corpus Persian Corpus (Merged) Dataset Summary Persian Corpus (Merged) is a large-scale, Persian corpus meticulously aggregated from multiple high-quality Persian datasets available on the Hugging Face Hub. Designed to advance Persian NLP research and applications, this corpus consolidates diverse textual sources into a single resource, providing researchers and developers with a robust foundation for training and evaluating language models. Why Use This… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-text-corpus.texttext-generation10M<n<100M0 likes434 downloads2mo agoHugging Face06PotatoHD /ru-text-corpus Description 798k deduplicated Russian documents (1.6B tokens) from FineWeb-2 (rus_Cyrl), ru-StackOverflow, Pikabu, Habr, Russian Wikipedia and news. Markdown- and code-bearing (StackOverflow/Habr/Pikabu are text_markdown). Filtered to >=200 chars and >=30% Cyrillic. Columns: text, src. Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is. Usage from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/PotatoHD/ru-text-corpus.texttext-generation100K<n<1M0 likes429 downloads4mo agoHugging Face07IRIIS-RESEARCH /Nepali-Text-Corpus Nepali Text Corpus Overview Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.texttext-generation1M<n<10M10 likes337 downloads1y agoHugging Face08mridul3301 /nepali-text-corpus-64 Nepali Text Dataset Overview The Nepali Text Dataset is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics. Dataset Details Total Articles: ~6.4 million Language:… See the full description on the dataset page: https://huggingface.co/datasets/mridul3301/nepali-text-corpus-64.text1M<n<10M5 likes264 downloads2y agoHugging Face09khtsly /roblox_docs_corpus_text [!Note] Last collected: 2026-10-01 15:11 Contains 1789 contents texttext-generation1K<n<10K1 likes168 downloads9d agoHugging Face10Baybars /parla_text_corpus ParlaTextCorpus Spoken text corpus for Catalan. Derived and cleaned from three sources. OpenSubtitles, Tv3Parla and Festcat. text100K<n<1M0 likes157 downloads4y agoHugging Face11PereLluis13 /parla_text_corpustext1M<n<10M0 likes141 downloads5y agoHugging Face12paodigitalhub /blk-text-corpus Verified Pa'O (Blk) Text Corpus - Pa'O Digital Hub Dataset Summary This is the official, verified parallel dataset for the Pa'O language (ISO 639-3: blk) and Burmese (Myanmar) translations, published by Pa'O Digital Hub. The corpus is systematically collected, reviewed, standardized, and verified through the established linguistic and editorial workflow of Pa'O Digital Hub. The Pa'O sentences are based on authentic language usage by Pa'O native speakers and are… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/blk-text-corpus.texttranslationn<1K1 likes132 downloads4d agoHugging Face13Panhapich /khmer-text-corpus Khmer + English Mixed Text Corpus A cleaned, deduplicated Khmer text corpus with naturally-occurring Khmer/English code-switching (English tech & finance terms, Latin script, and digits embedded in Khmer text). It is the training data for the Panhapich/khmer-sp-8k SentencePiece tokenizer and the text-only warm-start of a shared Khmer diffusion decoder. Dataset summary 4,893,739 sentences, one per line, UTF-8, deduplicated and shuffled (fixed seed 42… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-text-corpus.texttext-generation1M<n<10M2 likes128 downloads3mo agoHugging Face14kalixlouiis /burmese-text-corpus Burmese Text Corpus For Natural Language Processing 🎫 Choose your language: 🌏 English Version | 🇲🇲 မြန်မာဗားရှင်း 🌏 English Version This dataset is a specifically curated text corpus for the Burmese language. It is intended to support Natural Language Processing (NLP) tasks, language model training, and research related to the Burmese language. 1. About the Dataset The primary goal of creating this burmese-text-corpus dataset is to address the scarcity of… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/burmese-text-corpus.texttext-classification1K<n<10K12 likes110 downloads6mo agoHugging Face15MatinaAI /matina_persian_text_corpusgated Matina: A Large-Scale 73B Token Persian Text Corpus Text corpora are essential for training models used in tasks like summarization, translation, and large language models (LLMs). While various efforts have been made to collect monolingual and multilingual datasets in many languages, Persian has often been underrepresented due to limited resources for data collection and preprocessing. Existing Persian datasets are typically small and lack content diversity, consisting mainly of… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/matina_persian_text_corpus.texttext-generation11 likes108 downloads1y agoHugging Face16Pangeanic /Cantonese-Japanese-Machine-Translation-Corpus-text PangeanicYueJa - Cantonese Japanese Parallel Corpus PangeanicYueJa is a Cantonese-Japanese parallel corpus designed for machine translation, multilingual large language model (LLM) training, cross-lingual NLP research, retrieval-augmented generation (RAG), bilingual embeddings, instruction tuning, and multilingual AI systems. This release contains 55,000 Cantonese-Japanese sentence pairs sampled from a larger corpus of approximately 3.08 million parallel sentence pairs. For the… See the full description on the dataset page: https://huggingface.co/datasets/Pangeanic/Cantonese-Japanese-Machine-Translation-Corpus-text.texttranslation10K<n<100K8 likes107 downloads4mo agoHugging Face17dewozniak /video-text-corpus Code Video Text Data Notes Dataset summary A documented Code data-preparation workflow for Video Text records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus. Included material build_dataset.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small, human-readable records for checking the schema. README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/dewozniak/video-text-corpus.0 likes107 downloads29d agoHugging Face18thomaswardzij /video-text-corpus Movie Posters Video Text Data Notes Dataset summary A documented Movie Posters data-preparation workflow for Video Text records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus. Included material clean.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small, human-readable records for checking the schema.… See the full description on the dataset page: https://huggingface.co/datasets/thomaswardzij/video-text-corpus.0 likes89 downloads25d agoHugging Face19Oscarbaek /video-text-corpus Math Video Text Data Notes Dataset summary This repository contains a preparation pipeline and a small metadata sample for Math work with Video Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated. Included material dataloader.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small, human-readable… See the full description on the dataset page: https://huggingface.co/datasets/Oscarbaek/video-text-corpus.0 likes87 downloads26d agoHugging Face20subramanianno /video-text-corpus51 Comics Video Text Data Notes Dataset summary This repository contains a preparation pipeline and a small metadata sample for Comics work with Video Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated. Included material dataloader.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/subramanianno/video-text-corpus51.0 likes81 downloads25d agoHugging Face21Mwanzau /Tumbuka_Text_Corpus_Translated_Gutenberg Tumbuka Text Corpus - Translated Gutenberg Dataset Description This dataset contains a large-scale collection of Tumbuka text, primarily consisting of machine-translated literary works from the Project Gutenberg library. It is designed to support Natural Language Processing (NLP) research for Tumbuka, a Bantu language spoken in Malawi, Zambia, and Tanzania. Dataset Summary Language: Tumbuka (tum) Source: Project Gutenberg Content: Translated… See the full description on the dataset page: https://huggingface.co/datasets/Mwanzau/Tumbuka_Text_Corpus_Translated_Gutenberg.text100M<n<1B0 likes76 downloads3mo agoHugging Face22raygx /Nepali-Extended-Text-Corpus Dataset Card for "Nepali-Extended-Text-Corpus" More Information needed text10M<n<100M3 likes73 downloads3y agoHugging Face23heitorlr /text-tabular-corpus Gaming Text Tabular Data Notes Dataset summary Preparation notes and schema examples for Gaming tasks using Text Tabular data. Full source material is intentionally not bundled, so provenance and licensing remain explicit. Included material preprocess.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small, human-readable records for checking the schema. README.md — data card… See the full description on the dataset page: https://huggingface.co/datasets/heitorlr/text-tabular-corpus.0 likes72 downloads28d agoHugging Face24vietnqw /vi-text_corpus-dantri.com.vn-splittedtext100K<n<1M0 likes67 downloads2y agoHugging Face25FELIXFF92 /video-text-corpus Geology Video Text Data Notes Dataset summary Preparation notes and schema examples for Geology tasks using Video Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit. Included material preprocess.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small, human-readable records for checking the schema. README.md — data card… See the full description on the dataset page: https://huggingface.co/datasets/FELIXFF92/video-text-corpus.0 likes67 downloads8d agoHugging Face26jolawal1995 /image-text-corpus Speech Image Text Data Notes Dataset summary A documented Speech data-preparation workflow for Image Text records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus. Included material build_dataset.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small, human-readable records for checking the schema. README.md —… See the full description on the dataset page: https://huggingface.co/datasets/jolawal1995/image-text-corpus.0 likes66 downloads28d agoHugging Face27Adeptschneider /CiviVox-English-Swahili-text-translation-corpustextn<1K0 likes65 downloads2y agoHugging Face28Reubencf /adaption-amharic-text-corpus This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. adaption-amharic_text_corpus This dataset comprises over 700,000 Amharic text documents formatted as line-delimited JSON, covering diverse topics such as history, religion, politics, and product descriptions. Each entry contains a single string field with native Amharic content, including some samples with mixed languages or placeholder values. It is designed for text… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/adaption-amharic-text-corpus.text10K<n<100K0 likes64 downloads3mo agoHugging Face29TianfuXinqu /product_review_text_corpus Academic Product Review Text Corpus Academic corpus of product review texts assembled by a university research group. Terms Released under the Creative Commons Attribution 4.0 International license (CC-BY-4.0). 0 likes61 downloads2mo agoHugging Face30cody-jones /text-tabular-corpus Agriculture Text Tabular Data Notes Dataset summary This repository contains a preparation pipeline and a small metadata sample for Agriculture work with Text Tabular inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated. Included material prepare.py — loading, cleaning, and split preparation code. dataset_infos.json — schema and split metadata. metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/cody-jones/text-tabular-corpus.0 likes61 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.