datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fast_3_tasks_multilingual_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 121,
"total_frames": 31783,
"total_tasks": 9,
"total_videos": 242,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:121"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Ruszczka/fast_3_tasks_multilingual_v1.vatex-multilingual-retrieval
VATEX multilingual (en/zh) video retrieval (MTEB)
Cross-lingual video retrieval over VATEX, which
captions every clip in both English and Chinese. The two subsets share one video
corpus and differ only in caption language, which makes a like-for-like cross-lingual
comparison possible.
Prepared for MTEB as
VATEXMultilingualT2VRetrieval and VATEXMultilingualV2TRetrieval.
Contents
config
rows
description
videos
993
shared video corpus, 10s clips
en
993… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/vatex-multilingual-retrieval.test-groot-CoffeePressButton-hindi-spanish
