Team Ai
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01applied-ai-018 /peacock-data-public-datasets-idc-enfm-dataprocessing BharatGPT Data Curation for English Foundational Model RedPajamaV2 RedPajamaV2 directory contains scripts for processing and filtering of RedPajamaV2 dataset. Overall pipeline is as follows: URL Filtering scripts/utils/url_filtering.py: The first step of filtering is URL based. We use a blocklist of URLs known to contain inappropriate content. This list is taken from blocklistproject. It contains categories for advertisements, gambling, adult content, etc. All… See the full description on the dataset page: https://huggingface.co/datasets/applied-ai-018/peacock-data-public-datasets-idc-enfm-dataprocessing.0 likes274 downloads2y agoHugging Face02renjiepi /easy_5000_data_processingtext10K<n<100K0 likes99 downloads9mo agoHugging Face03renjiepi /medium_5000_data_processingtext1K<n<10K0 likes83 downloads9mo agoHugging Face04laion /nemotron-terminal-data_processing nemotron-terminal-data_processing Per-source partition of nvidia/Nemotron-Terminal-Corpus, filtered to source == "data_processing". The difficulty column preserves the original easy / medium / mixed split (na for the dataset_adapters/* files, which did not carry a difficulty label). Partitioning scheme: adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet {skill} (e.g. debugging, security, …) — rows from synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-data_processing.textquestion-answering1K<n<10K0 likes68 downloads6mo agoHugging Face05renjiepi /medium_5000-data_processing_n100k1text1K<n<10K0 likes61 downloads9mo agoHugging Face06renjiepi /medium_5000_data_processing_fixedtext1K<n<10K0 likes60 downloads9mo agoHugging Face07Nexdata-kr /1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample Description 한국어 시험 문제 구조화 분석·가공 데이터로, 약 150만 개의 시험 문제를 포함하고 있습니다. 문제 유형, 문제, 정답, 해설 등의 정보를 포함하며, 과목은 [초등학교] 국어, 수학, 영어, 사회, 과학; [중학교] 국어, 영어, 수학, 과학, 사회; [고등학교] 국어, 영어, 수학, 물리, 화학, 생물, 역사, 지리로 구성되어 있습니다. 문제 유형에는 객관식, 빈칸 채우기, 참·거짓 문제, 단답형 문제 등이 포함됩니다. 본 데이터셋은 대규모 교과 지식 강화 및 학습 데이터 구축 등의 작업에 활용할 수 있습니다. 자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/1634?source=Hf.kr Specifications Data content 한국어 K12 시험 문제 Amount 약 150만 개의… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.textn<1K0 likes52 downloads1mo agoHugging Face08renjiepi /easy_5000-data_processing_n100k1text1K<n<10K0 likes51 downloads9mo agoHugging Face09uv-scripts /data-processing Data processing Part of uv-scripts — self-contained UV scripts you run on Hugging Face Jobs in one command. General data processing recipes: convert, clean and prepare data files. Script What it does optimize-parquet.py Converts CSV, JSON and Parquet files uploaded to a bucket into optimized Parquet, triggered by a bucket webhook optimize-parquet.py: optimized Parquet from bucket uploads Upload a CSV, JSON or Parquet file to a Storage Bucket and… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/data-processing.0 likes47 downloads12d agoHugging Face10renjiepi /mixed_1000_data_processingtextn<1K0 likes43 downloads9mo agoHugging Face11renjiepi /easy_5000_data_processing_fixedtext1K<n<10K0 likes38 downloads9mo agoHugging Face12renjiepi /mixed_1000-data_processing_1000_n30k1textn<1K0 likes37 downloads9mo agoHugging Face13DCAgent2 /terminal_bench_2_nemotron_terminal_data_processing__Qwen3_8B_20260413_170737textn<1K0 likes36 downloads6mo agoHugging Face14renjiepi /mixed_1000_data_processing_fixedtextn<1K0 likes28 downloads9mo agoHugging Face15Nexdata-AI /1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample Description Korean Test Questions Structured Analysis Processing Data, around 1.5 million questions, contains question types, questions, answers, explanations, etc..For subjects, include [Primary School] Korean, Mathematics, English, Social Studies, Science; [Middle School] Korean, English, Mathematics, Science, Social Studies; [High School] Korean, English, Mathematics, Physics, Chemistry, Biology, History, Geography; question Types indlude single-choice question, fill-in… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.textn<1K0 likes22 downloads2mo agoHugging Face16Sraghvi /data_processing_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "dual_arm_robot", "total_episodes": 1, "total_frames": 3105, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Sraghvi/data_processing_test.tabularrobotics1K<n<10K0 likes20 downloads1y agoHugging Face17snowdere /autotrain-data-stratefied-processing AutoTrain Dataset for project: stratefied-processing Dataset Description This dataset has been automatically processed by AutoTrain for project stratefied-processing. Languages The BCP-47 code for the dataset's language is unk. Dataset Structure Data Instances A sample from this dataset looks as follows: [ { "text": "dateline resources ltd", "target": 0 }, { "text": "dateline resources", "target": 0 } ]… See the full description on the dataset page: https://huggingface.co/datasets/snowdere/autotrain-data-stratefied-processing.text-classification0 likes18 downloads4y agoHugging Face18rombodawg /data_processing_codeThese are the scripts I used to clean the rombodawg/code_bagel and rombodawg/code_bagel_hermes-2.5 datasets. In order for these scripts to work your datasets need to be in the format bellow, with the condition that each line is its own .json object. {"instruction": "", "input": "", "output": ""} {"instruction": "", "input": "", "output": ""} {"instruction": "", "input": "", "output": ""} {"instruction": "", "input": "", "output": ""} {"instruction": "", "input": "", "output": ""} 4 likes18 downloads2y agoHugging Face19DCAgent2 /swebench_verified_random_100_folders_nemotron_terminal_data_processing__Qwen3_87dd0272etextn<1K0 likes18 downloads6mo agoHugging Face20nahiar /sentiment_data_train_id_en_sentiment_30k_post-processingtext10K<n<100K0 likes14 downloads7mo agoHugging Face21DCAgent2 /dev_set_v2_nemotron_terminal_data_processing__Qwen3_8B_20260413_175806textn<1K0 likes13 downloads6mo agoHugging Face22nahiar /sentiment_30k_data_train_sentimen_id_post-processingtext10K<n<100K0 likes8 downloads7mo agoHugging Face23114-Digital-Image-Processing-Group18 /splat-dataimagen<1K0 likes6 downloads10mo agoHugging Face24nahiar /sentiment_data_train_sentimen_id_post-processingtext10K<n<100K0 likes5 downloads7mo agoHugging Face25Nexdata-AI /10-million-English-Test-Questions-Text-Parsing-And-Processing-Data-Sample Description 10 Million - English Test Questions Text Parsing And Processing Data, Each question contains title, answer, parse, subject, grade, question type; The educational stages cover primary, middle, high school, and university; Subjects cover mathmatics, biology, accounting, etc.The data are questions text under the Anglo-American system, which can be used to enhance the subject knowledge of large models For more details, please refer to the link:… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/10-million-English-Test-Questions-Text-Parsing-And-Processing-Data-Sample.textn<1K0 likes5 downloads2mo agoHugging Face26OsakanaTeishoku /structured_data_with_cot_dataset_512_v2_dpo_before_processingtext1K<n<10K0 likes4 downloads9mo agoHugging Face27nahiar /spam_post-processing-data-train-final-25k-id-entext10K<n<100K0 likes4 downloads7mo agoHugging Face28infra777 /my-data-processingtext1K<n<10K0 likes4 downloads3mo agoHugging Face29gametelounge /video-processing-data0 likes3 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.