Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01VLABench /raw_primitive_datasets Datacard This is the official fine-tuning dataset provided by VLABench (raw data), with 500 episodes each task. The current version includes 10 primitive tasks. Source Project Page: https://vlabench.github.io/ Arxiv Paper: https://arxiv.org/abs/2412.18194 Code: https://github.com/OpenMOSS/VLABench Uses Download all archive files and use the following command to extract: cat vlabench_primitive.tar.gz.* | tar -xzvf - In the resulting… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/raw_primitive_datasets.text1K<n<10K4 likes2.1k downloads6mo agoHugging Face02Santhosh1884 /IXI-Datasetstext1K<n<10K0 likes543 downloads10mo agoHugging Face03OpenMOSS-Team /FRoM-W1-Datasets FRoM-W1: Towards General Humanoid Whole-Body Control with Language Instructions The Humanoid Intelligence Team from FudanNLP and OpenMOSS Introduction Humanoid robots are capable of performing various actions such as greeting, dancing and even backflipping. However, these motions are often hard-coded or specifically trained, which limits their versatility. In this work, we present FRoM-W1[^1]… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FRoM-W1-Datasets.3d100K<n<1M8 likes472 downloads8mo agoHugging Face04CASIA-LM /OpenS2S_Datasets How to Use? Download, merge the files, and extract You can run the following command to merge the compressed file parts after downloading. cat en_response_wav.tar.gz.* > en_response_wav.tar.gz cat zh_response_wav.tar.gz.* > zh_response_wav.tar.gz audio100K<n<1M8 likes449 downloads1y agoHugging Face05hqfang /sam2act-datasets SAM2Act SAM2Act is a multi-view robotics transformer policy for robotic manipulation. Built on RVT-2, it combines multi-resolution upsampling with visual embeddings from the SAM2 foundation model to improve 3D action prediction, multitask learning, and generalization. SAM2Act+ extends this policy with a memory bank, memory encoder, and memory attention so the agent can condition on prior observations and actions for spatial memory-dependent tasks. For full project details, code… See the full description on the dataset page: https://huggingface.co/datasets/hqfang/sam2act-datasets.textrobotics100K<n<1M2 likes367 downloads5mo agoHugging Face06qihoo360 /VRF-datasetsimage100K<n<1M0 likes263 downloads6mo agoHugging Face07gsarch /vigorl_datasets ViGoRL Datasets This repository contains the official datasets associated with the paper "Grounded Reinforcement Learning for Visual Reasoning (ViGoRL)", by Gabriel Sarch, Snigdha Saha, Naitik Khandelwal, Ayush Jain, Michael J. Tarr, Aviral Kumar, and Katerina Fragkiadaki. Dataset Overview These datasets are designed for training and evaluating visually grounded vision-language models (VLMs). Datasets are organized by the visual reasoning tasks described in the ViGoRL… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/vigorl_datasets.imagevisual-question-answering100K<n<1M1 likes246 downloads1y agoHugging Face08zeyuren2002 /AnyDepth_datasetstextn<1K0 likes170 downloads5mo agoHugging Face09gate-institute /GATE-VLAP-datasets GATE-VLAP Datasets Grounded Action Trajectory Embeddings with Vision-Language Action Planning This repository contains preprocessed datasets from the LIBERO benchmark suite in WebDataset TAR format, specifically designed for training vision-language-action models with semantic action segmentation. Data Format: WebDataset TAR We provide datasets in WebDataset TAR format for optimal performance: ✅ Fast loading - Efficient streaming during training✅ Easy downloading - Single… See the full description on the dataset page: https://huggingface.co/datasets/gate-institute/GATE-VLAP-datasets.imagereinforcement-learning100K<n<1M3 likes157 downloads10mo agoHugging Face10hackroot /Wan_datasets rCM: Score-Regularized Continuous-Time Consistency Model Paper | Website | Code This repo holds Wan-synthesized datasets used for rCM training. Citation @article{zheng2025rcm, title={Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency}, author={Zheng, Kaiwen and Wang, Yuji and Ma, Qianli and Chen, Huayu and Zhang, Jintao and Balaji, Yogesh and Chen, Jianfei and Liu, Ming-Yu and Zhu, Jun and Zhang, Qinsheng}, journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/hackroot/Wan_datasets.text100K<n<1M0 likes83 downloads10mo agoHugging Face11wifibk /CFNet_Datasetsimage100K<n<1M1 likes77 downloads2y agoHugging Face12hjvsl /GeoZero_Train_Datasets SFT and RL Traning dataset of GeoZero Dataset Composition GeoZero consists of three variants: File Description GeoZero-Raw.json Raw aggregated data across heterogeneous datasets GeoZero-Instruct.json Unified instruction-tuned dataset for supervised fine-tuning GeoZero-Hard.json Challenging subset for RL training All image files are stored under the images/ directory. Directory Structure GeoZero_Train_Datasets/ ├── images/ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/hjvsl/GeoZero_Train_Datasets.imagevisual-question-answering100K<n<1M2 likes63 downloads8mo agoHugging Face13Duceh /datasetsimage10K<n<100K0 likes51 downloads2y agoHugging Face14talas /pwm_datasetstext100K<n<1M0 likes31 downloads11mo agoHugging Face15keshavagarwal2004dev /GeoZero_Train_Datasets SFT and RL Traning dataset of GeoZero Dataset Composition GeoZero consists of three variants: File Description GeoZero-Raw.json Raw aggregated data across heterogeneous datasets GeoZero-Instruct.json Unified instruction-tuned dataset for supervised fine-tuning GeoZero-Hard.json Challenging subset for RL training All image files are stored under the images/ directory. Directory Structure GeoZero_Train_Datasets/ ├── images/ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/keshavagarwal2004dev/GeoZero_Train_Datasets.imagevisual-question-answering100K<n<1M0 likes25 downloads7mo agoHugging Face16TTS-AGI /balanced-audio-score-datasets-DACVAEtext100K<n<1M0 likes22 downloads7mo agoHugging Face17gsarch /vigorl_gaze_datasetstext1K<n<10K0 likes18 downloads1y agoHugging Face18thomascsh /SOC-defection-datasetstext10K<n<100K0 likes17 downloads5mo agoHugging Face19jvarga92 /patchman_datasetstext100K<n<1M0 likes14 downloads1y agoHugging Face20zhangyupeng90 /labelgs_datasetsimage1K<n<10K1 likes14 downloads1y agoHugging Face21ATG2222 /my-gec-datasetstextn<1K0 likes14 downloads6mo agoHugging Face22furman-lab /patchman_datasetstext10K<n<100K0 likes11 downloads1y agoHugging Face23lucasgblu /nanostep-datasetsimagen<1K0 likes9 downloads4y agoHugging Face24xDAN-datasets /img-demoimagen<1K0 likes9 downloads2y agoHugging Face25datasets-examples /doc-audio-11 [doc] audio dataset 11 This dataset contains two tar files that contain pairs of samples with one audio file and one JSON file. audion<1K0 likes8 downloads2y agoHugging Face26Mohinta2892 /Catena_datasetsimagen<1K0 likes8 downloads9mo agoHugging Face27wliu283 /datasets_49text1M<n<10M0 likes8 downloads4mo agoHugging Face28Tobivictor /legal_datasets Legal Case Documents Dataset Snapshot for GaiaNet This repository contains a Qdrant database snapshot of vectorized legal case documents from Nigeria and the UK. The data has been processed and converted into embeddings suitable for use as a knowledge base in a Retrieval-Augmented Generation (RAG) system. Purpose This dataset was created to serve as a specialized knowledge base for the GaiaNet ecosystem. The goal is to enable a Large Language Model (LLM) to answer… See the full description on the dataset page: https://huggingface.co/datasets/Tobivictor/legal_datasets.textn<1K0 likes4 downloads1y agoHugging Face29RacheinWu /NL-To1_sft_datasetsimage1K<n<10K0 likes4 downloads4mo agoHugging Face30chunping-hf /dataset_splitgatedaudio10K<n<100K0 likes2 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.