Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RekaAI /RekaDaily-10k-processed RekaDaily-10k (processed) Short first-person clips cut from the RekaDaily-10k recordings — unscripted daily-life video collected through Claru, Reka's data collection marketplace, recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Every clip carries one dense caption and a multi-question Q&A exchange written in the second person ("What am I doing in this video?"), so the corpus drops straight into… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-processed.imagevideo-text-to-text1M<n<10M4 likes76k downloads28d agoHugging Face02zackyabd /ptb-xl-processedtabular10K<n<100K0 likes73k downloads1y agoHugging Face03jkot /parliament_hearings_processed Preprocessed parliament hearings ASR dataset to truecased form. Original dataset: https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3126 dataset_info: features: - name: id dtype: string - name: audio dtype: audio: sampling_rate: 16000 - name: transcription sequence: string splits: - name: train num_bytes: 53645064353.18 num_examples: 191455 - name: test num_bytes: 740331298.0 num_examples: 2726… See the full description on the dataset page: https://huggingface.co/datasets/jkot/parliament_hearings_processed.audio100K<n<1M1 likes16k downloads3y agoHugging Face04garak-llm /drh-System-Prompt-processedtextn<1K0 likes11k downloads6mo agoHugging Face05KMK040412 /aitw-processed-labeled-full AiTW Processed Full with App Labels This repository contains a full processed Android in the Wild (AiTW) mirror together with an app-labeled step index, official split assignment by episode_id, major-app statistics, and a ready-to-train Gmail subset. Why This Exists AiTW is large and not easy to navigate by app. The original labels contain useful fields such as goal_info, current_activity, and action coordinates, but users often need extra processing before they… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/aitw-processed-labeled-full.imageimage-text-to-text1M<n<10M1 likes8.5k downloads4mo agoHugging Face06BrunoHays /multilingual_librispeech_fr_processed multilingual_librispeech_fr_processed Dataset Description Dataset Summary The data files can be found on the illuin gcloud instance at this adress: unknown_url This dataset has been processed from Huggingface Hub dataset facebook/multilingual_librispeech and the config french Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/multilingual_librispeech_fr_processed.text100K<n<1M1 likes7.3k downloads2y agoHugging Face07open-r1 /DAPO-Math-17k-Processed Dataset Card for DAPO-Math-17k-Processed This is a processed version of BytedTsinghua-SIA/DAPO-Math-17k where we have: Deduplicated the prompts Reformatted the prompts and ground truth answers to be compatible with TRL's GRPO trainer We have also derived pure English and Chinese subsets. The full dataset processing logic can be found in create_dataset.py. If you find this dataset useful in your work, please cite the original source with: @misc{yu2025dapoopensourcellmreinforcement… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/DAPO-Math-17k-Processed.text10K<n<100K88 likes5.8k downloads11mo agoHugging Face08fjd /scannet-processed-testimage1 likes5.4k downloads4y agoHugging Face09gijs /avqa-processedaudio10K<n<100K0 likes3.9k downloads1y agoHugging Face10Qwen /ProcessBench ProcessBench This repository contains the dataset of the ProcessBench benchmark proposed by Qwen Team. You can refer to our GitHub repository for the evaluation code and the prompt templates we use in this work. If you find this work relevant or helpful to your work, please kindly cite us: @article{processbench, title={ProcessBench: Identifying Process Errors in Mathematical Reasoning}, author={ Chujie Zheng and Zhenru Zhang and Beichen Zhang and Runji Lin and Keming Lu and… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/ProcessBench.text1K<n<10K59 likes3.7k downloads2y agoHugging Face11noxwano /ASMR-Archive-Processed-SFW ASMR-Archive-Processed-SFW Overview This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset. We filtered the original dataset to include only records where the nsfw metadata flag is false. To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled. The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.audioautomatic-speech-recognition1M<n<10M9 likes3.5k downloads6mo agoHugging Face12wandb /RAGTruth-processed RAGTruth Dataset Dataset Description Dataset Summary The RAGTruth dataset is designed for evaluating hallucinations in text generation models, particularly in retrieval-augmented generation (RAG) contexts. It contains examples of model outputs along with expert annotations indicating whether the outputs contain hallucinations. Dataset Structure Each example contains: A query/question Context passages Model output Hallucination labels (evident… See the full description on the dataset page: https://huggingface.co/datasets/wandb/RAGTruth-processed.text10K<n<100K31 likes3k downloads2y agoHugging Face13Tri1 /processed_vnhntabular10K<n<100K0 likes2.4k downloads2mo agoHugging Face14ptllama /processed_acemath_fulltext1M<n<10M0 likes2k downloads2y agoHugging Face15Xuhui /sft_processed_large_split sft_processed_large — profile-disjoint split This is the train / val / test split of Xuhui/sft_processed_large, the OdysSim midtraining corpus (21.4M interactions across 63 datasets). Split structure split rows how it's built train 21.20M what's left after val + test are carved out val 28K per-dataset random sample, in-distribution; for checkpoint selection test 128K profile-disjoint where the dataset's profile space supports it; for OOD generalization… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sft_processed_large_split.texttext-generation10M<n<100M1 likes1.4k downloads6mo agoHugging Face16krahets /dna_rendering_processedgated DNA-Rendering-Processed Dataset Project Page | Paper | Code | Model To enable Diffuman4D model training, we meticulously process the DNA-Rendering dataset by recalibrating camera parameters, optimizing image color correction matrices (CCMs), predicting foreground masks, and estimating human skeletons. To promote future research in the field of human-centric 3D/4D generation, we have open-sourced our re-annotated labels for the DNA-Rendering dataset in this repo, which includes… See the full description on the dataset page: https://huggingface.co/datasets/krahets/dna_rendering_processed.imageimage-to-3d1K<n<10K9 likes1.3k downloads11mo agoHugging Face17typesafe /evalsafe-invoice-processing Invoice processing Snapshot: 2026-09-28. 150 cases and 6,874 question instances. Default reference: consensus. Labels are model-generated references. Data Load configuration cases, questions, or run_results; all have a test split. cases: one row per case_id, with the complete input in input_json, descriptive metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped by policy_id and contain status, actions, and primary_action. questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-invoice-processing.tabular1K<n<10K6 likes1.3k downloads11d agoHugging Face18vikp /doclaynet_processed Dataset Card for "doclaynet_processed" Clean version of DocLayNet ready for finetuning. image10K<n<100K6 likes1.2k downloads3y agoHugging Face19wandb /finqa-data-processed FinQA Dataset (Processed) Dataset Description Dataset Summary The FinQA dataset is designed for numerical reasoning over financial data, containing questions that require complex reasoning over tables and text from financial reports. Dataset Statistics Total examples: 8281 Training set size: 6624 examples Test set size: 1657 examples Dataset Structure Each example contains: Required columns: query: The question to be answered (derived… See the full description on the dataset page: https://huggingface.co/datasets/wandb/finqa-data-processed.text1K<n<10K2 likes1.1k downloads2y agoHugging Face20henryscheible /coco_val2014_blip2_processed Dataset Card for "coco_val2014_blip2_processed" More Information needed text10K<n<100K0 likes1.1k downloads4y agoHugging Face21ayz2 /drivaernet_processed DrivAerNet++ (processed) This is a processed, downsampled version of DrivAerNet++, not the original dataset. Fields were converted to a common frame and non-dimensionalized, rows were randomly subsampled and some variables were dropped. For the original data, see the original paper The original size around 2.3TB, therefore the volume fields was downsampled 5x, and the surface field is kept as is to reduce the size to 0.78 TB. Layout collated/… See the full description on the dataset page: https://huggingface.co/datasets/ayz2/drivaernet_processed.text10K<n<100K0 likes956 downloads3d agoHugging Face222084Collective /deepstock-stock-historical-prices-dataset-processedtabular10M<n<100M0 likes887 downloads2y agoHugging Face23tuanmanh28 /VIVOS_CommonVoice_FOSD_Control_processed_dataset Dataset Card for "VIVOS_CommonVoice_FOSD_Control_processed_dataset" More Information needed audio10K<n<100K2 likes838 downloads3y agoHugging Face24Jakh0103 /glotlid_processedtext100M<n<1B1 likes814 downloads2y agoHugging Face25LennardZuendorf /openlegaldata-processed Dataset Card for openlegaldata.io bulk case data Dataset Description This is a edit/cleanup of Bulk Data of openlegaldata.io, which I also brought onto Huggingface here. The Entire Dataset Is In German Github Repository: [uniArchive-legalis]](https://github.com/LennardZuendorf/uniArchive-legalis) Repository: Bulk Data Edit Summary I have done some cleaning and splitting of the data and filtered out large parts that were not (easily) usable, cutting… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/openlegaldata-processed.texttext-classification1K<n<10K2 likes808 downloads3y agoHugging Face26hitachi-nlp /proofwriter_processed_OWAtabular10K<n<100K2 likes791 downloads2y agoHugging Face27MihailSlutsky /vistr-process-verification-pilot ViSTR Process-Verification Pilot (14 answer-correct trajectories, multimodal) Agent trajectories for studying process false positives in multimodal agents: cases where the answer is correct but the visual reasoning that produced it is wrong. Ships the raw perception tool outputs so any claim in a trajectory can be independently re-verified, plus human annotations and an unmodified XSkill critique of the same trajectories. Why this exists Harness / skill… See the full description on the dataset page: https://huggingface.co/datasets/MihailSlutsky/vistr-process-verification-pilot.imagevisual-question-answeringn<1K0 likes732 downloads2mo agoHugging Face28dgorbatov /vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10 Trajectory Ranking Dataset This dataset contains trajectory ranking results for autonomous navigation scenarios. Dataset Statistics Total examples: 39558 Chunks processed: 40 Upload date: 2025-09-13T00:44:30.335177 Features Image data with terrain analysis Trajectory rankings and reasoning Quality and diversity analysis Terrain and trajectory descriptions imageimage-classification10K<n<100K0 likes708 downloads1y agoHugging Face29himalaya-ai /ocr-document-processing-eval ocr_document_processing_eval Document-processing OCR proxy for digitization, KYC-like numeric fields, and RAG extraction checks. Repo: himalaya-ai/ocr-document-processing-eval Task: document_processing_ocr Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns. Optional fine-tuning/eval file: *.sharegpt.json with messages and images. Core Columns id: unique sample identifier image: relative path to the image file ocr:… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/ocr-document-processing-eval.imageimage-to-textn<1K1 likes676 downloads4mo agoHugging Face30ZzZZCHS /processed_scannetimage100K<n<1M0 likes657 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.