datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
peacock-data-public-datasets-idc-enfm-dataprocessing
BharatGPT
Data Curation for English Foundational Model
RedPajamaV2
RedPajamaV2 directory contains scripts for processing and filtering of RedPajamaV2 dataset.
Overall pipeline is as follows:
URL Filtering scripts/utils/url_filtering.py: The first step of filtering is URL based. We use a blocklist of URLs known to contain inappropriate content. This list is taken from blocklistproject. It contains categories for advertisements, gambling, adult content, etc. All… See the full description on the dataset page: https://huggingface.co/datasets/applied-ai-018/peacock-data-public-datasets-idc-enfm-dataprocessing.easy_5000_data_processingmedium_5000_data_processingnemotron-terminal-data_processing
nemotron-terminal-data_processing
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "data_processing". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-data_processing.medium_5000-data_processing_n100k1medium_5000_data_processing_fixed1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample
Description
한국어 시험 문제 구조화 분석·가공 데이터로, 약 150만 개의 시험 문제를 포함하고 있습니다. 문제 유형, 문제, 정답, 해설 등의 정보를 포함하며, 과목은 [초등학교] 국어, 수학, 영어, 사회, 과학; [중학교] 국어, 영어, 수학, 과학, 사회; [고등학교] 국어, 영어, 수학, 물리, 화학, 생물, 역사, 지리로 구성되어 있습니다. 문제 유형에는 객관식, 빈칸 채우기, 참·거짓 문제, 단답형 문제 등이 포함됩니다. 본 데이터셋은 대규모 교과 지식 강화 및 학습 데이터 구축 등의 작업에 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/1634?source=Hf.kr
Specifications
Data content
한국어 K12 시험 문제
Amount
약 150만 개의… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.easy_5000-data_processing_n100k1data-processing
Data processing
Part of uv-scripts — self-contained UV scripts you run on Hugging Face Jobs in one command.
General data processing recipes: convert, clean and prepare data files.
Script
What it does
optimize-parquet.py
Converts CSV, JSON and Parquet files uploaded to a bucket into optimized Parquet, triggered by a bucket webhook
optimize-parquet.py: optimized Parquet from bucket uploads
Upload a CSV, JSON or Parquet file to a Storage Bucket and… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/data-processing.mixed_1000_data_processingeasy_5000_data_processing_fixedmixed_1000-data_processing_1000_n30k1terminal_bench_2_nemotron_terminal_data_processing__Qwen3_8B_20260413_170737mixed_1000_data_processing_fixed1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample
Description
Korean Test Questions Structured Analysis Processing Data, around 1.5 million questions, contains question types, questions, answers, explanations, etc..For subjects, include [Primary School] Korean, Mathematics, English, Social Studies, Science; [Middle School] Korean, English, Mathematics, Science, Social Studies; [High School] Korean, English, Mathematics, Physics, Chemistry, Biology, History, Geography; question Types indlude single-choice question, fill-in… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.data_processing_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "dual_arm_robot",
"total_episodes": 1,
"total_frames": 3105,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Sraghvi/data_processing_test.autotrain-data-stratefied-processing
AutoTrain Dataset for project: stratefied-processing
Dataset Description
This dataset has been automatically processed by AutoTrain for project stratefied-processing.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": "dateline resources ltd",
"target": 0
},
{
"text": "dateline resources",
"target": 0
}
]… See the full description on the dataset page: https://huggingface.co/datasets/snowdere/autotrain-data-stratefied-processing.data_processing_codeThese are the scripts I used to clean the rombodawg/code_bagel and rombodawg/code_bagel_hermes-2.5 datasets.
In order for these scripts to work your datasets need to be in the format bellow, with the condition that each line is its own .json object.
{"instruction": "", "input": "", "output": ""}
{"instruction": "", "input": "", "output": ""}
{"instruction": "", "input": "", "output": ""}
{"instruction": "", "input": "", "output": ""}
{"instruction": "", "input": "", "output": ""}
swebench_verified_random_100_folders_nemotron_terminal_data_processing__Qwen3_87dd0272esentiment_data_train_id_en_sentiment_30k_post-processingdev_set_v2_nemotron_terminal_data_processing__Qwen3_8B_20260413_175806sentiment_30k_data_train_sentimen_id_post-processingsplat-datasentiment_data_train_sentimen_id_post-processing10-million-English-Test-Questions-Text-Parsing-And-Processing-Data-Sample
Description
10 Million - English Test Questions Text Parsing And Processing Data, Each question contains title, answer, parse, subject, grade, question type; The educational stages cover primary, middle, high school, and university; Subjects cover mathmatics, biology, accounting, etc.The data are questions text under the Anglo-American system, which can be used to enhance the subject knowledge of large models
For more details, please refer to the link:… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/10-million-English-Test-Questions-Text-Parsing-And-Processing-Data-Sample.structured_data_with_cot_dataset_512_v2_dpo_before_processingspam_post-processing-data-train-final-25k-id-enmy-data-processingvideo-processing-data
