Team Ai
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01paulpacaud /rlbenchfail_test_dataset Guardian: RLBench-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.tabularvisual-question-answering1K<n<10K0 likes892 downloads8mo agoHugging Face02paulpacaud /ur5fail_test_dataset Guardian Failure Detection Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_test_dataset.tabularvisual-question-answeringn<1K1 likes203 downloads8mo agoHugging Face03paulpacaud /bdv2fail_test_dataset Guardian: BridgeDataV2-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_test_dataset.tabularvisual-question-answering1K<n<10K0 likes104 downloads8mo agoHugging Face04danfperam /testdatatabular10K<n<100K0 likes40 downloads1y agoHugging Face05YatharthDedhia /TestDatasetScratch repo for testing dataset-viewer schema inference. All data is synthetic and contains no real personal information. case_bank/ holds a few raw per-case JSON files whose nested mock block is union-typed across tools (some lists are empty, some hold structs), which breaks single-schema inference. viewer/cases.jsonl is a flattened table with a stable schema, and the configs block above points the viewer at it so the raw files are not globbed. tabularn<1K0 likes39 downloads1mo agoHugging Face06dlthub /test_refresh_staging_dataset74a5fe5b763e5e0e93259b2b1e0af7d4_data_stagingtabularn<1K0 likes29 downloads4d agoHugging Face07jhyun0414 /3_0507_test_dataset_final 3_0507_test_dataset_final This dataset contains medical test questions and answers. Files data.json: Array of objects with keys like question_id, question, choices, correct_answer, etc. tabular1K<n<10K0 likes17 downloads1y agoHugging Face08data4elm /roleplay-test roleplay-test Test split for lm-eval-harness. tabular10K<n<100K0 likes16 downloads1y agoHugging Face09semran1 /test_datatabular10K<n<100K0 likes15 downloads8mo agoHugging Face10timqian /test-datasettabularn<1K0 likes14 downloads1y agoHugging Face11dlthub /test_data_202608010526003995_stagingtabularn<1K0 likes12 downloads2mo agoHugging Face12rescommons /Ecom-Chatbot-Synthetic-Test-Datasettabular1K<n<10K0 likes9 downloads7mo agoHugging Face13Schweinhund /schweinhunds_test_datasettabularn<1K0 likes6 downloads4y agoHugging Face14bingqin111 /scored_test_data_by_rm_filtertabular10K<n<100K0 likes6 downloads1y agoHugging Face15yhc2222 /tsp_testdatatabularn<1K0 likes6 downloads7mo agoHugging Face16Schandkroete /English_Skills_Test_DatasetgatedThis is our testing dataset for the skills. It contains all skill groups of the skills in our query samples. tabular1K<n<10K0 likes5 downloads3y agoHugging Face17zhaospei /data-50k-test-finaltabular1K<n<10K0 likes5 downloads3y agoHugging Face18nishitneema /MBPP-cluster_0-based-fewshot-prompting-test-dataset-take2tabularn<1K0 likes5 downloads2y agoHugging Face19hugochien /testdataset2tabularn<1K0 likes3 downloads3y agoHugging Face20zhaospei /data-50k-refine-test-gen2tabular1K<n<10K0 likes3 downloads3y agoHugging Face21WhoCares10 /TestDatasettabularn<1K0 likes3 downloads2y agoHugging Face22nishitneema /MBPP-cluster-based-fewshot-prompting-test-datasettabularn<1K0 likes3 downloads2y agoHugging Face23dlthub /schema_test_data_202602191021032182tabularn<1K0 likes3 downloads8mo agoHugging Face24snehangshuk /test-datasettabular1K<n<10K0 likes2 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.