Team Ai
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tvu-vlinhd11 /pretrain-dataset-raw-10M Pretrain Dataset (Text) This dataset contains preprocessed text documents ready for LLM pretraining. Dataset Details Property Value Documents 10,000,000 Processed 10000000 Shards 21 Created 2025-12-09 Dataset Structure Each sample contains: text: The document text source: Source dataset identifier id: Unique document ID Usage from datasets import load_dataset dataset = load_dataset("tvu-vlinhd11/pretrain-dataset-raw-10M")… See the full description on the dataset page: https://huggingface.co/datasets/tvu-vlinhd11/pretrain-dataset-raw-10M.texttext-generation10M<n<100M0 likes514 downloads10mo agoHugging Face02daitavan /Vietnam-Law-Raw-Datatext-generation1 likes261 downloads2y agoHugging Face03AmanPriyanshu /rlvr-guru-raw-data-extended RLVR GURU Extended: Compiling a 150K Cross-Domain Dataset for RLVR A comprehensive cross-domain reasoning dataset containing 150,000 training samples and 221,332 test samples across diverse reasoning-intensive domains. This dataset extends the foundational work from the GURU dataset (Cheng et al., 2025) by incorporating additional STEM reasoning domains (MedMCQA and CommonsenseQA) while maintaining rigorous quality standards and verification mechanisms essential for reinforcement… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/rlvr-guru-raw-data-extended.texttext-generation100K<n<1M0 likes260 downloads1y agoHugging Face04AgenticFinLab /multiagent-entropy-rawdata When Does Multi-Agent Collaboration Help? An Entropy Perspective 📄 Paper · 💻 Code · 🌐 Project Page This is the raw experimental data behind every figure and claim in paper. Use it to reproduce all results and conclusions reported in the paper. Data Overview Size: ~5 GB, 237 filesFormat: CSV (aggregated metrics) and JSON (entropy distributions, evaluation metrics) The data is organized as follows: 1. Merged Dataset merged_datasets/master.csv… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/multiagent-entropy-rawdata.text-generation8 likes179 downloads4mo agoHugging Face05BEE-spoke-data /govdocs1-txt-raw Dataset Card for "govdocs1-txt-raw" Somewhere to put the raw txt files before filtering them Source info/page: https://digitalcorpora.org/corpora/file-corpora/files/ @inproceedings{garfinkel2009bringing, title={Bringing Science to Digital Forensics with Standardized Forensic Corpora}, author={Garfinkel, Simson and Farrell, Paul and Roussev, Vassil and Dinolt, George}, booktitle={Digital Forensic Research Workshop (DFRWS) 2009}, year={2009}, address={Montreal, Canada}… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-txt-raw.texttext-generation10K<n<100K0 likes145 downloads10mo agoHugging Face06snupilab /aka-llama-korean-dataset-multiturn-raw Aka-LLAMA Korean Multi-Turn Dataset (Raw) This dataset is a raw version of a multi-turn Korean conversation dataset generated using kordinal. It is designed for research and development in Korean natural language processing (NLP), specifically in multi-turn dialogue generation. License This dataset is released under the CC BY-NC 4.0 license. It is strictly for non-commercial research and educational purposes. Commercial usage is prohibited. Additionally, some data… See the full description on the dataset page: https://huggingface.co/datasets/snupilab/aka-llama-korean-dataset-multiturn-raw.textquestion-answering10K<n<100K3 likes97 downloads2y agoHugging Face07Roshan32 /Hinglish_Dataset_instruction_and_rawtexttext-generation10K<n<100K1 likes96 downloads9mo agoHugging Face08BEE-spoke-data /napierone-epub-raw BEE-spoke-data/napierone-epub-raw NapierOne EPUB files converted with marker. Seems to contain mostly books from Project Gutenberg. detected languages via fasttext-langdetect {'ca': 1, 'cy': 1, 'da': 6, 'de': 105, 'en': 4403, 'eo': 2, 'es': 61, 'fi': 76, 'fr': 189, 'he': 1, 'hu': 5, 'is': 1, 'it': 40, 'la': 6, 'nl': 41, 'pl': 4, 'pt': 38, 'sv': 10, 'tl': 9} texttext-generation10K<n<100K0 likes86 downloads10mo agoHugging Face09freococo /myawady-raw-dataset Myawady Raw News Corpus 🇲🇲 This dataset contains over 59,000 full-text Burmese news articles scraped from the Myawady News Portal, the official media outlet of the Myanmar military government. Unlike the title-only version, this dataset includes complete article content, with metadata fields such as category, publication date, and image URLs. It is intended for use in Myanmar NLP and AI research, including: 🧠 Language modeling 📰 Text summarization 🏷️ Named entity… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myawady-raw-dataset.imagetext-classification10K<n<100K0 likes37 downloads1y agoHugging Face10kalixlouiis /raw-data Dataset Card for kalixlouiis/raw-data This dataset is a high-quality, large-scale Burmese language corpus curated from a diverse range of sources, including classical literature, news, encyclopedic content, and conversational data. It has been rigorously cleaned to ensure linguistic integrity and is suitable for various NLP tasks such as Language Modeling, Tokenizer Training, and Text Classification. Dataset Summary The kalixlouiis/raw-data corpus is a synthesized… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/raw-data.textfeature-extraction1M<n<10M5 likes37 downloads8mo agoHugging Face11datalamhnyin /myanmar-raw Myanmar Raw Sentences Corpus (datalamhnyin/myanmar-raw) A collection of Burmese text sentences compiled from various public sources, intended for Myanmar NLP and language model pre-training research. Status (Raw/Uncleaned):This dataset is in its raw stage. Basic deduplication and length filtering have been applied, but it may contain OCR noise, non-standard spelling variations, or encoding artifacts. Further task-specific normalization is recommended before deployment.… See the full description on the dataset page: https://huggingface.co/datasets/datalamhnyin/myanmar-raw.texttext-generation1M<n<10M0 likes34 downloads7d agoHugging Face12BEE-spoke-data /napierone-pdf-raw BEE-spoke-data/napierone-pdf-raw NapierOne PDF files converted with marker. detected languages Counter({'en': 4665, 'nl': 2, 'fi': 7, 'fr': 8, 'cy': 54, 'sq': 1, 'it': 1, 'unknown-error': 5, 'sk': 1, 'es': 2, 'de': 3, 'ro': 1, 'pl': 1, 'zh': 1, 'so': 1, 'ml': 1}) tabulartext-generation10K<n<100K0 likes25 downloads10mo agoHugging Face13dippatel2506 /agri-llm-raw-datasetgated Agri-LLM Raw Dataset Dataset Description The Agri-LLM Raw Dataset is a collection of processed text extracted from various PDF documents related to agriculture. This dataset is intended for use in language modeling and other NLP tasks focused on agricultural content. Dataset Structure Data Fields sentence_chunks: A list of sentence chunks, where each chunk contains up to 10 sentences. Example Data Below is an example of a single entry in… See the full description on the dataset page: https://huggingface.co/datasets/dippatel2506/agri-llm-raw-dataset.texttext-generation1K<n<10K5 likes13 downloads2y agoHugging Face14nickting /nyt-connections-datasets-raw NYT Connections Raw Datasets This repository contains the raw and formatted reasoning data for NYT Connections puzzle solving experiments. These files are the source data used to create the experiment splits in nickting/nyt-connections-experiments. Overview This dataset includes three types of puzzle data with AI-generated reasoning: NYT Connections Puzzles - Authentic New York Times puzzles Synthetic Connections Puzzles - Algorithmically generated puzzles… See the full description on the dataset page: https://huggingface.co/datasets/nickting/nyt-connections-datasets-raw.question-answering10K<n<100K0 likes10 downloads1y agoHugging Face15SilvioLima /raw_data textfeature-extraction10K<n<100K0 likes4 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.