AV
Models
All models matching “AV”Datasets
All datasets matching “AV”avm_residential_dataMomentSeekerMomentSeeker: A Comprehensive Benchmark and A Strong Baseline For Moment Retrieval Within Long Videos
This repo contains the annotation data for the paper "MomentSeeker: A Comprehensive Benchmark and A Strong Baseline For Moment Retrieval Within Long Videos".
🔔 News:
🥳 2025/03/07: We have released the MomentSeeker Benchmark and Paper! 🔥
Introduction
We present MomentSeeker, a comprehensive benchmark to… See the full description on the dataset page: https://huggingface.co/datasets/avery00/MomentSeeker.Physical-AI-AV-US
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 2,789,773 samples from 150 000
driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the
United States.
Format
WebDataset — 100 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.png
Front-facing wide-angle camera frame (640 × 360 px)
{key}.json
Metadata (see schema below)… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.smollm-corpus
SmolLM-Corpus: Now shuffled and sharded!
This is a version of the SmolLM-Corpus where the 3 subsets have been interleved, shuffled and sharded as 23698 jsonl.zst files for easy streaming!
The dataset is comprised of the cosmopedia-v2 and fineweb-edu-dedup subsets from the original SmolLM-Corpus repo, with the python-edu subset being pulled from my python-edu repo.
Dataset Structure
The dataset is split into 24 subdirectories, with the first 23 containing 1000 shards… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/smollm-corpus.avspeech-visual-audio
AVSpeech Video + Audio
This repository is a media-bearing reconstruction of the public AVSpeech
annotations. Each row represents an already-trimmed segment and keeps the
original source-video timing and target-face-center metadata.
Dataset structure
clip_id: identifier derived as
{youtube_id}_{start_sec:.3f}_{end_sec:.3f}.
avspeech_metadata: JSON containing youtube_id, start_sec, end_sec,
x_center, and y_center from the AVSpeech annotation.
video: video-only… See the full description on the dataset page: https://huggingface.co/datasets/ProgramComputer/avspeech-visual-audio.AVQA-videos
AVQA — Audio-Visual Question Answering (videos + annotations)
A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life
audio-visual question answering over short in-the-wild clips. The original release
ships only the QA annotations and expects users to collect the source videos from
VGGSound themselves. This repository bundles the source video clips together
with the official train/val annotations, so the dataset is usable without any
YouTube scraping.… See the full description on the dataset page: https://huggingface.co/datasets/juyil/AVQA-videos.
