Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RekaAI /RekaDaily-10k-processed RekaDaily-10k (processed) Short first-person clips cut from the RekaDaily-10k recordings — unscripted daily-life video collected through Claru, Reka's data collection marketplace, recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Every clip carries one dense caption and a multi-question Q&A exchange written in the second person ("What am I doing in this video?"), so the corpus drops straight into… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-processed.imagevideo-text-to-text1M<n<10M4 likes75k downloads27d agoHugging Face02zackyabd /ptb-xl-processedtabular10K<n<100K0 likes73k downloads1y agoHugging Face03meihualuomanxueshan /Processed_Interiorverse0 likes20k downloads2y agoHugging Face04jkot /parliament_hearings_processed Preprocessed parliament hearings ASR dataset to truecased form. Original dataset: https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3126 dataset_info: features: - name: id dtype: string - name: audio dtype: audio: sampling_rate: 16000 - name: transcription sequence: string splits: - name: train num_bytes: 53645064353.18 num_examples: 191455 - name: test num_bytes: 740331298.0 num_examples: 2726… See the full description on the dataset page: https://huggingface.co/datasets/jkot/parliament_hearings_processed.audio100K<n<1M1 likes17k downloads3y agoHugging Face05robometer /processed_datasets2 likes13k downloads9mo agoHugging Face06vctvct123 /BlendedMVS_processedimage100K<n<1M0 likes12k downloads6mo agoHugging Face07garak-llm /drh-System-Prompt-processedtextn<1K0 likes11k downloads6mo agoHugging Face08brownu /deform360_processed1 likes9k downloads2mo agoHugging Face09KMK040412 /aitw-processed-labeled-full AiTW Processed Full with App Labels This repository contains a full processed Android in the Wild (AiTW) mirror together with an app-labeled step index, official split assignment by episode_id, major-app statistics, and a ready-to-train Gmail subset. Why This Exists AiTW is large and not easy to navigate by app. The original labels contain useful fields such as goal_info, current_activity, and action coordinates, but users often need extra processing before they… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/aitw-processed-labeled-full.imageimage-text-to-text1M<n<10M1 likes8.3k downloads4mo agoHugging Face10meihualuomanxueshan /Processed_interiorverse_851 likes8.2k downloads2y agoHugging Face11BrunoHays /multilingual_librispeech_fr_processed multilingual_librispeech_fr_processed Dataset Description Dataset Summary The data files can be found on the illuin gcloud instance at this adress: unknown_url This dataset has been processed from Huggingface Hub dataset facebook/multilingual_librispeech and the config french Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/multilingual_librispeech_fr_processed.text100K<n<1M1 likes7.4k downloads2y agoHugging Face12data-process /QVHighlights-zip0 likes7k downloads1y agoHugging Face13open-r1 /DAPO-Math-17k-Processed Dataset Card for DAPO-Math-17k-Processed This is a processed version of BytedTsinghua-SIA/DAPO-Math-17k where we have: Deduplicated the prompts Reformatted the prompts and ground truth answers to be compatible with TRL's GRPO trainer We have also derived pure English and Chinese subsets. The full dataset processing logic can be found in create_dataset.py. If you find this dataset useful in your work, please cite the original source with: @misc{yu2025dapoopensourcellmreinforcement… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/DAPO-Math-17k-Processed.text10K<n<100K88 likes6k downloads11mo agoHugging Face14fjd /scannet-processed-testimage1 likes5.4k downloads4y agoHugging Face15omrastogi /Hypersim-Processedimage0 likes4.9k downloads2y agoHugging Face16gijs /avqa-processedaudio10K<n<100K0 likes4k downloads1y agoHugging Face17Qwen /ProcessBench ProcessBench This repository contains the dataset of the ProcessBench benchmark proposed by Qwen Team. You can refer to our GitHub repository for the evaluation code and the prompt templates we use in this work. If you find this work relevant or helpful to your work, please kindly cite us: @article{processbench, title={ProcessBench: Identifying Process Errors in Mathematical Reasoning}, author={ Chujie Zheng and Zhenru Zhang and Beichen Zhang and Runji Lin and Keming Lu and… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/ProcessBench.text1K<n<10K59 likes3.8k downloads2y agoHugging Face18infinity1096 /flyingthings3d_processed0 likes3.6k downloads1y agoHugging Face19sanagnos /processed_gpt_dataset_big Dataset Card for "processed_gpt_dataset_big" More Information needed 1M<n<10M0 likes2.9k downloads4y agoHugging Face20wandb /RAGTruth-processed RAGTruth Dataset Dataset Description Dataset Summary The RAGTruth dataset is designed for evaluating hallucinations in text generation models, particularly in retrieval-augmented generation (RAG) contexts. It contains examples of model outputs along with expert annotations indicating whether the outputs contain hallucinations. Dataset Structure Each example contains: A query/question Context passages Model output Hallucination labels (evident… See the full description on the dataset page: https://huggingface.co/datasets/wandb/RAGTruth-processed.text10K<n<100K30 likes2.9k downloads2y agoHugging Face21noxwano /ASMR-Archive-Processed-SFW ASMR-Archive-Processed-SFW Overview This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset. We filtered the original dataset to include only records where the nsfw metadata flag is false. To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled. The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.audioautomatic-speech-recognition1M<n<10M9 likes2.9k downloads6mo agoHugging Face22OmniAICreator /ASMR-Archive-Processed ASMR-Archive-Processed (WIP) Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended. Work in Progress — expect breaking changes while the pipeline and data layout stabilize. This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.imageautomatic-speech-recognition98 likes2.4k downloads6mo agoHugging Face23Tri1 /processed_vnhntabular10K<n<100K0 likes2.4k downloads2mo agoHugging Face24Fanqi-Lin /Processed-Task-Dataset Robotic Manipulation Datasets for Four Tasks [Project Page] [Paper] [Code] [Models] [Raw GoPro Videos] This repository contains in-the-wild robotic manipulation datasets collected using UMI, and processed through a SLAM pipeline, as described in the paper "Data Scaling Laws in Imitation Learning for Robotic Manipulation". The datasets cover four tasks: Pour Water Arrange Mouse Fold Towel Unplug Charger Dataset Folders: arrange_mouse and pour_water: Each folder contains… See the full description on the dataset page: https://huggingface.co/datasets/Fanqi-Lin/Processed-Task-Dataset.robotics100B<n<1T9 likes2.4k downloads2y agoHugging Face25ssmits /processed-falcon-dutch-dataset100K<n<1M0 likes2.3k downloads2y agoHugging Face26LLM-OS-Models /Qwen-Terminal-ToolBench-Processed-Tokenized Qwen Terminal ToolBench Processed Datasets Qwen-family processed/template-applied and selected tokenized terminal datasets. Contents qwen35_2b_full_terminal_toolcall_processed_v1: 1011776 rows, 36.26 GiB, features: text qwen35_2b_terminal_toolcall_fullconv_from_samecount_v1: 327383 rows, 25.11 GiB, features: text qwen35_27b_hf_fsdp_1024: 15335 rows, 0.17 GiB, features: input_ids, attention_mask, labels qwen35_27b_processed_2bdata: 23472 rows, 2.77 GiB, features:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/Qwen-Terminal-ToolBench-Processed-Tokenized.text-generation0 likes2.2k downloads4mo agoHugging Face27Kavindu1124 /ucf-crime-processed-features0 likes2.2k downloads1mo agoHugging Face28ptllama /processed_acemath_fulltext1M<n<10M0 likes2k downloads2y agoHugging Face29nevvton /subgraphrag-processed-emb0 likes1.9k downloads2h agoHugging Face30Ken4962 /processed_fake_job_postingstabulartext-classification10K<n<100K0 likes1.8k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.