ProgramComputer/avspeech-visual-audio
AVSpeech Video + Audio This repository is a media-bearing reconstruction of the public AVSpeech annotations. Each row represents an already-trimmed segment and keeps the original source-video timing and target-face-center metadata. Dataset structure clip_id: identifier derived as {youtube_id}_{start_sec:.3f}_{end_sec:.3f}. avspeech_metadata: JSON containing youtube_id, start_sec, end_sec, x_center, and y_center from the AVSpeech annotation. video: video-only… See the full description on the dataset page: https://huggingface.co/datasets/ProgramComputer/avspeech-visual-audio.
58.2k
1---2configs:3- config_name: default4 data_files:5 - split: train6 path: train/*.parquet7 - split: test8 path: test/*.parquet9---10 11# AVSpeech Video + Audio12 13This repository is a media-bearing reconstruction of the public AVSpeech14annotations. Each row represents an already-trimmed segment and keeps the15original source-video timing and target-face-center metadata.16 17## Dataset structure18 19- `clip_id`: identifier derived as20 `{youtube_id}_{start_sec:.3f}_{end_sec:.3f}`.21- `avspeech_metadata`: JSON containing `youtube_id`, `start_sec`, `end_sec`,22 `x_center`, and `y_center` from the AVSpeech annotation.23- `video`: video-only stream, or null when the source segment could not be24 materialized.25- `audio`: audio-only stream, or null when the source segment could not be26 materialized.27 28The `video` field is already trimmed. `start_sec` and `end_sec` refer to the29original YouTube-video timeline and must not be used to seek again within this30clip. AVSpeech defines `(x_center, y_center)` as the normalized center of the31speaker's face in the frame at the beginning of the segment, with `(0, 0)` at32the top left.33 34## Audited snapshot and known publication gap35 36The following figures describe revision37`efdceb2a0b9d81a6aec76f10668cca49e8209e37`:38 39| Split | Rows | Rows with both media | Rows without both media | Parquet files | Encoded size |40| --- | ---: | ---: | ---: | ---: | ---: |41| train | 2,621,845 | 1,589,842 | 1,032,003 | 5,142 | 1,404,473,032,806 bytes |42| test | 183,273 | 98,605 | 84,668 | 359 | 88,861,772,006 bytes |43| total | 2,805,118 | 1,688,447 | 1,116,671 | 5,501 | 1,493,334,804,812 bytes |44 45The train states were verified by an exhaustive read-only scan of all 5,14246train Parquet files and 2,621,845 rows. Of the 1,032,003 train rows without both47media streams, 1,032,000 have both media null, 3 are video-only, and 0 are48audio-only. The test split was not audited at that row-state granularity, so49the table reports only its aggregate count without both streams.50 51The completed exporter expected 1,589,942 paired train occurrences, while this52snapshot contains 1,589,842, an aggregate gap of 100. Surviving non-media53evidence supports high-confidence assignment of 49 of those occurrence slots:54 55- 46 both-null occurrences have a paired sibling at the pinned revision.56- 3 video-only occurrences retain a published `video.path` while `audio.path`57 is null, showing that the exporter reached archive-member processing.58 59The remaining 51 occurrence slots cannot be assigned to exact rows or60`clip_id` values without the original expected-pair manifest or historical61ID-to-archive map. Their ambiguity remains within a pool of 4,10362metadata-resolved both-null occurrences across 403 YouTube IDs. “Unresolved”63does not mean that these rows were verified unavailable.64 65The audit selected only `clip_id`, `avspeech_metadata`, `video.path`, and66`audio.path`. It did not select or materialize embedded media bytes, and its67range guards recorded zero intersections with media-byte column chunks. No68media recovery, YouTube retrieval, torrent-media transfer, row repair, or69Hugging Face mutation was attempted as part of that audit.70 71## Bounded loading72 73Do not use `snapshot_download` for routine training: the repository is about741.49 TB. Stream rows, keep media decoding disabled at the dataset layer, skip75any row without both media streams, and materialize only one bounded work unit76at a time. Preserve the official split and stable row provenance, and report77pre-filter and retained denominators plus exclusions by reason.78 79```python80from datasets import Audio, Video, load_dataset81 82revision = "efdceb2a0b9d81a6aec76f10668cca49e8209e37"83rows = load_dataset(84 "ProgramComputer/avspeech-visual-audio",85 split="train",86 revision=revision,87 streaming=True,88)89rows = rows.cast_column("video", Video(decode=False))90rows = rows.cast_column("audio", Audio(decode=False))91 92for row in rows:93 if row["video"] is None or row["audio"] is None:94 continue95 # Materialize/process this row in bounded temporary storage.96```97 98The official AVSpeech page states that its supplied train and test annotations99use disjoint speakers. This reconstruction preserves those source split labels.100It does not add person identities, and a YouTube video ID must not be described101as a speaker identity.102 103## Intended use and limitations104 105This dataset is intended for research on audio-visual speech and related106representation-learning tasks. It is derived from public Internet video and is107not demographically balanced. Availability, codecs, media quality, language,108pose, lighting, and annotation accuracy vary. Missing rows are not necessarily109random, so filtering to paired media may introduce additional selection bias.110The unresolved 51-slot publication gap is aggregate provenance information,111not a verified unavailable-row list, and must not be converted into invented112row-level labels.113 114The face-center coordinate is a point hint at the beginning of the segment,115not a bounding box, persistent track, verified identity label, or consent116signal. Downstream systems must validate the associated detected face and must117not use this dataset for identification, surveillance, or consequential118decisions.119 120## License and provenance review121 122The official AVSpeech download page provides train/test annotation CSVs and123states that “this data” is available under CC BY 4.0. This repository also124redistributes media derived from YouTube videos. The maintainer has not yet125documented a legal review establishing that the same license statement covers126redistribution of every embedded media stream or that all upstream platform127and uploader terms are satisfied. Therefore this card deliberately does not128assert a Hugging Face `license` tag for the media-bearing reconstruction.129 130Before continued public redistribution, document the source acquisition131process, takedown procedure, upstream terms, and the basis for redistributing132the embedded audio/video. This note is a publication safeguard, not legal133advice.134 135## Citation136 137If you use the data, cite the original AVSpeech work:138 139```bibtex140@article{ephrat2018looking,141 title={Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation},142 author={Ephrat, Ariel and Mosseri, Inbar and Lang, Oran and Dekel, Tali and Wilson, Kevin and Hassidim, Avinatan and Freeman, William T. and Rubinstein, Michael},143 journal={ACM Transactions on Graphics},144 year={2018}145}146```147 148Official AVSpeech project and download page:149<https://looking-to-listen.github.io/avspeech/download.html>.150 