datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
testdatanew_dataset_testtest_raw_video_data
ShareGPTVideo Raw Videos for Testing data
All dataset and models can be found at ShareGPTVideo.
Contents:
In case of need, this contains raw videos corresponding to test frames in
Test video frames
DicFace-test_datasettest-dataTest_dataset_repoRobin-test-dataThis dataset contains the first 40k prompts from LAION/CC/SBU BLIP-Caption Concept-balanced 558K which we use for rapid testing of the Robin model setup on new compute.
This is based on the data used in LLaVA: https://github.com/haotian-liu/LLaVA/blob/main/docs/Data.md
This does about 150 iterations with a batch size of 256 to check checkpointing and final model save.
test_hf_data_5
test_hf_data_5
This is a video dataset in WebDataset format.
Structure
Each .tar archive contains video files (.mp4) and metadata (.json) with the same prefix.
The JSON files contain:
file_name
label (numeric)
categories (emotion category string)
description (human annotation)
How to use
from datasets import load_dataset
ds = load_dataset("ZebangCheng/test_hf_data_5", split="train", streaming=True)
sample = next(iter(ds))
# Save video
with… See the full description on the dataset page: https://huggingface.co/datasets/ZebangCheng/test_hf_data_5.vox-vietnam-test-dataTest-DATAk_test_datasettest-dataset-alexmy-dataset-test-v2test_datatestdata
