datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bambara-whisper-featuresmajestrino-1.00-16xk5-sae-features
Majestrino 1.00 SAE — Feature Audio Samples (16x, k=5)
Top-2000 activating audio samples for each feature in the
Majestrino 1.00 SAE.
Overview
Metric
Value
SAE Architecture
16x expansion, k=5, d_model=768
Total Features
12,288
Alive Features
10,684
Audio per Feature
Up to 2,000 highest-activating
Audio Format
Opus (24 kbps OGG container)
Total TAR Files
1069
Source Dataset
laion/majestrino-data
File Structure
Each TAR file… See the full description on the dataset page: https://huggingface.co/datasets/laion/majestrino-1.00-16xk5-sae-features.TCGA_foundation_model_featuresgalaxy_video_clip_featuresRACnet_feature_npy
RACnet_feature_npy
This dataset is the *_feature_npz folders for RACnet
Dataset Details
❗Attention: without annotation files
qvhighlight_internvideo2_llama_text_featureimagenet1k_features_256_sd_vae_ft_emaCASTELLA_CLAP_features
CASTELLA CLAP features
This repository contains audio and text features of CASTELLA dataset extracted by CLAP.
Using these features, we can reproduce the audio moments retrieval using CASTELLA, which is used in lighthouse.
Please also check demo page.
How to Download?
Run the following script:
from huggingface_hub import snapshot_download
repo_id = "lighthouse-emnlp2024/CASTELLA_CLAP_features"
local_dir = "./"
downloaded_path = snapshot_download(… See the full description on the dataset page: https://huggingface.co/datasets/lighthouse-emnlp2024/CASTELLA_CLAP_features.ego4d_videomae_L14_feature_fps8
📙 Overview
Ego4d video features extracted by VideoMAE_L14 at 8 fps.
It contains 9645 files, each file (e.g. fffbaeef-577f-45f0-baa9-f10cabf62dfb.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768.
🏋️ How-To-Use
Please refer to code EgoInstructor for details.
🎓 Citation
@article{xu2024retrieval,
title={Retrieval-augmented egocentric video captioning},
author={Xu, Jilan and Huang, Yifei and Hou, Junlin and Chen, Guo… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/ego4d_videomae_L14_feature_fps8.imagenet_features_1024_sd_vae_ft_emaObjaverseXL_sketchfab-features-part_0002charade_sta_internvideo2_llama_text_featureObjaverseXL_sketchfab-features-part_0004egolearn_videomae_internvideo_features
📙 Overview
Egolearn video features.
egocentric videos are extracted by VideoMAE_L14 at 8 fps.
exocentric videos are extracted by InternVideo_MM_L14 at 8 fps.
They are used for evaluating the egovideo-exovideo retrieval ability of EgoInstructor.
Each file (e.g. 2cad4224-56c4-11ee-88ee-80615f12b59e.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768.
🏋️ How-To-Use
Please refer to code EgoInstructor for details.
🎓… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/egolearn_videomae_internvideo_features.youcook2_internvideo_MM_L14_features_fps8
📙 Overview
YouCook2 video features extracted by InternVideo_MM_L14 at 8 fps. It is used for evaluating the video-text retrieval ability of EgoInstructor.
Each file (e.g. 10dZTHlkb8w.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768.
🏋️ How-To-Use
Please refer to code EgoInstructor for details.
🎓 Citation
@article{xu2024retrieval,
title={Retrieval-augmented egocentric video captioning},
author={Xu, Jilan and Huang… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/youcook2_internvideo_MM_L14_features_fps8.charadesego_videomae_L14_feature_fps8
📙 Overview
CharadesEgo video features extracted by VideoMAE_L14 at 8 fps. It is used for evaluating the egovideo-exovideo retrieval ability of EgoInstructor.
It contains 7860 files, each file (e.g. 005BUEGO.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768.
🏋️ How-To-Use
Please refer to code EgoInstructor for details.
🎓 Citation
@article{xu2024retrieval,
title={Retrieval-augmented egocentric video captioning}… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/charadesego_videomae_L14_feature_fps8.microvent-features
microvent-features
Derived signals for the microvent core release: per-keyframe OCR text,
per-chunk ASR transcripts, and an embedding zoo (keyframe-level vision,
keyframe-OCR text, audio-level, video-level, omni-modal).
This card covers only the features. For the source videos, audio,
keyframes, and the public eval annotations, see the microvent dataset
card. All artifacts here key on the same chunk_id and follow the same
WebDataset shard layout, so joining feature shards back… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/microvent-features.epic_kitchen_videomae_L14_feature_fps8
📙 Overview
EPIC-Kitchen-100 video features extracted by VideoMAE_L14 at 8 fps. It is used for evaluating the video-text retrieval ability of EgoInstructor.
It contains 700 files, each file (e.g. P01_01.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768.
🏋️ How-To-Use
Please refer to code EgoInstructor for details.
🎓 Citation
@article{xu2024retrieval,
title={Retrieval-augmented egocentric video captioning},
author={Xu… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/epic_kitchen_videomae_L14_feature_fps8.charadesego_internvideo_MM_L14_features_fps8
📙 Overview
CharadesEgo video features extracted by InternVideo_MM_L14 at 8 fps. It is used for evaluating the egovideo-exovideo retrieval ability of EgoInstructor.
It contains 7860 files, each file (e.g. 005BU.pth.tar) is a TxD feature vector, where T refers to the length of the video and D is 768.
🏋️ How-To-Use
Please refer to code EgoInstructor for details.
🎓 Citation
@article{xu2024retrieval,
title={Retrieval-augmented egocentric video captioning}… See the full description on the dataset page: https://huggingface.co/datasets/Jazzcharles/charadesego_internvideo_MM_L14_features_fps8.sd_vae_features_imagenet_1k_256x256recon_featurereconstruction_featurewasmedge-features
