datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pretrain-dataset-raw-10M
Pretrain Dataset (Text)
This dataset contains preprocessed text documents ready for LLM pretraining.
Dataset Details
Property
Value
Documents
10,000,000
Processed
10000000
Shards
21
Created
2025-12-09
Dataset Structure
Each sample contains:
text: The document text
source: Source dataset identifier
id: Unique document ID
Usage
from datasets import load_dataset
dataset = load_dataset("tvu-vlinhd11/pretrain-dataset-raw-10M")… See the full description on the dataset page: https://huggingface.co/datasets/tvu-vlinhd11/pretrain-dataset-raw-10M.Vietnam-Law-Raw-Datarlvr-guru-raw-data-extended
RLVR GURU Extended: Compiling a 150K Cross-Domain Dataset for RLVR
A comprehensive cross-domain reasoning dataset containing 150,000 training samples and 221,332 test samples across diverse reasoning-intensive domains. This dataset extends the foundational work from the GURU dataset (Cheng et al., 2025) by incorporating additional STEM reasoning domains (MedMCQA and CommonsenseQA) while maintaining rigorous quality standards and verification mechanisms essential for reinforcement… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/rlvr-guru-raw-data-extended.multiagent-entropy-rawdata
When Does Multi-Agent Collaboration Help? An Entropy Perspective
📄 Paper · 💻 Code · 🌐 Project Page
This is the raw experimental data behind every figure and claim in paper. Use it to reproduce all results and conclusions reported in the paper.
Data Overview
Size: ~5 GB, 237 filesFormat: CSV (aggregated metrics) and JSON (entropy distributions, evaluation metrics)
The data is organized as follows:
1. Merged Dataset
merged_datasets/master.csv… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/multiagent-entropy-rawdata.govdocs1-txt-raw
Dataset Card for "govdocs1-txt-raw"
Somewhere to put the raw txt files before filtering them
Source info/page: https://digitalcorpora.org/corpora/file-corpora/files/
@inproceedings{garfinkel2009bringing,
title={Bringing Science to Digital Forensics with Standardized Forensic Corpora},
author={Garfinkel, Simson and Farrell, Paul and Roussev, Vassil and Dinolt, George},
booktitle={Digital Forensic Research Workshop (DFRWS) 2009},
year={2009},
address={Montreal, Canada}… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-txt-raw.aka-llama-korean-dataset-multiturn-raw
Aka-LLAMA Korean Multi-Turn Dataset (Raw)
This dataset is a raw version of a multi-turn Korean conversation dataset generated using kordinal. It is designed for research and development in Korean natural language processing (NLP), specifically in multi-turn dialogue generation.
License
This dataset is released under the CC BY-NC 4.0 license. It is strictly for non-commercial research and educational purposes. Commercial usage is prohibited.
Additionally, some data… See the full description on the dataset page: https://huggingface.co/datasets/snupilab/aka-llama-korean-dataset-multiturn-raw.Hinglish_Dataset_instruction_and_rawnapierone-epub-raw
BEE-spoke-data/napierone-epub-raw
NapierOne EPUB files converted with marker. Seems to contain mostly books from Project Gutenberg.
detected languages
via fasttext-langdetect
{'ca': 1,
'cy': 1,
'da': 6,
'de': 105,
'en': 4403,
'eo': 2,
'es': 61,
'fi': 76,
'fr': 189,
'he': 1,
'hu': 5,
'is': 1,
'it': 40,
'la': 6,
'nl': 41,
'pl': 4,
'pt': 38,
'sv': 10,
'tl': 9}
myawady-raw-dataset
Myawady Raw News Corpus 🇲🇲
This dataset contains over 59,000 full-text Burmese news articles scraped from the Myawady News Portal, the official media outlet of the Myanmar military government.
Unlike the title-only version, this dataset includes complete article content, with metadata fields such as category, publication date, and image URLs. It is intended for use in Myanmar NLP and AI research, including:
🧠 Language modeling
📰 Text summarization
🏷️ Named entity… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myawady-raw-dataset.raw-data
Dataset Card for kalixlouiis/raw-data
This dataset is a high-quality, large-scale Burmese language corpus curated from a diverse range of sources, including classical literature, news, encyclopedic content, and conversational data. It has been rigorously cleaned to ensure linguistic integrity and is suitable for various NLP tasks such as Language Modeling, Tokenizer Training, and Text Classification.
Dataset Summary
The kalixlouiis/raw-data corpus is a synthesized… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/raw-data.myanmar-raw
Myanmar Raw Sentences Corpus (datalamhnyin/myanmar-raw)
A collection of Burmese text sentences compiled from various public sources, intended for Myanmar NLP and language model pre-training research.
Status (Raw/Uncleaned):This dataset is in its raw stage. Basic deduplication and length filtering have been applied, but it may contain OCR noise, non-standard spelling variations, or encoding artifacts. Further task-specific normalization is recommended before deployment.… See the full description on the dataset page: https://huggingface.co/datasets/datalamhnyin/myanmar-raw.napierone-pdf-raw
BEE-spoke-data/napierone-pdf-raw
NapierOne PDF files converted with marker.
detected languages
Counter({'en': 4665,
'nl': 2,
'fi': 7,
'fr': 8,
'cy': 54,
'sq': 1,
'it': 1,
'unknown-error': 5,
'sk': 1,
'es': 2,
'de': 3,
'ro': 1,
'pl': 1,
'zh': 1,
'so': 1,
'ml': 1})
agri-llm-raw-dataset
Agri-LLM Raw Dataset
Dataset Description
The Agri-LLM Raw Dataset is a collection of processed text extracted from various PDF documents related to agriculture. This dataset is intended for use in language modeling and other NLP tasks focused on agricultural content.
Dataset Structure
Data Fields
sentence_chunks: A list of sentence chunks, where each chunk contains up to 10 sentences.
Example Data
Below is an example of a single entry in… See the full description on the dataset page: https://huggingface.co/datasets/dippatel2506/agri-llm-raw-dataset.nyt-connections-datasets-raw
NYT Connections Raw Datasets
This repository contains the raw and formatted reasoning data for NYT Connections puzzle solving experiments. These files are the source data used to create the experiment splits in nickting/nyt-connections-experiments.
Overview
This dataset includes three types of puzzle data with AI-generated reasoning:
NYT Connections Puzzles - Authentic New York Times puzzles
Synthetic Connections Puzzles - Algorithmically generated puzzles… See the full description on the dataset page: https://huggingface.co/datasets/nickting/nyt-connections-datasets-raw.raw_data
