CVML-TueAI/Breakfast-Actions
๐ณ Breakfast Actions Dataset (HF + WebDataset Ready) This repository hosts the Breakfast Actions dataset metadata and videos, organized for modern deep learning workflows.It provides: 4 evaluation splits (s1, s2, s3, s4) JSONL metadata describing each video, participant, camera, and frame-level action segments Raw AVI videos stored directly on HuggingFace Optional WebDataset shards for streaming training ๐ Folder Layout Breakfast-Actions/ โ โโโโฆ See the full description on the dataset page: https://huggingface.co/datasets/CVML-TueAI/Breakfast-Actions.
๐ณ Breakfast Actions Dataset (HF + WebDataset Ready)
This repository hosts the Breakfast Actions dataset metadata and videos, organized for modern deep learning workflows. It provides:
- 4 evaluation splits (
s1,s2,s3,s4) - JSONL metadata describing each video, participant, camera, and frame-level action segments
- Raw AVI videos stored directly on HuggingFace
- Optional WebDataset shards for streaming training
๐ Folder Layout
Breakfast-Actions/
โ
โโโ Converted_Data/
โ โโโ metadata_s1.jsonl
โ โโโ metadata_s2.jsonl
โ โโโ metadata_s3.jsonl
โ โโโ metadata_s4.jsonl
โ
โโโ Videos/
โ โโโ P03/cam01/*.avi
โ โโโ P03/cam02/*.avi
โ โโโ P04/cam01/*.avi
โ โโโ ... (participants P03โP54, multiple cameras)
โ
โโโ WebDataset_Shards/ (optional)
โโโ 000000.tar
โโโ 000001.tar
โโโ ...๐ JSONL Record Format
Each metadata line looks like:
{
"video_path": "Videos/P03/cam01/P03_coffee.avi",
"participant": "P03",
"camera": "cam01",
"video": "P03_coffee",
"labels": [
{"start": 1, "end": 385, "label": "SIL"},
{"start": 385, "end": 599, "label": "pour_oil"},
...
]
}All video paths match the directory structure inside the HF repo.
๐น Load Metadata Using HuggingFace Datasets
from datasets import load_dataset
ds = load_dataset("json", data_files="metadata_s2.jsonl")["train"]
# Select all videos belonging to split s2
subset = ds๐น Load and Decode a Video
Using Decord
from decord import VideoReader
item = ds[0]
vr = VideoReader(item["video_path"])
frame0 = vr[0] # first frameUsing TorchVision
from torchvision.io import read_video
video, audio, info = read_video(item["video_path"])๐น WebDataset Version (Optional)
If the dataset includes .tar shards:
import webdataset as wds, jsonlines
ids = [rec["video_path"] for rec in jsonlines.open("metadata_s2.jsonl")]
dset = wds.WebDataset("WebDataset_Shards/*.tar").select(lambda s: s["json"]["video_path"] in ids)Each shard contains:
xxx.aviโ video bytesxxx.jsonโ metadata JSON
๐น PyTorch Example
from torch.utils.data import Dataset, DataLoader
from decord import VideoReader
class BreakfastDataset(Dataset):
def __init__(self, subset): self.subset = subset
def __len__(self): return len(self.subset)
def __getitem__(self, idx):
item = self.subset[idx]
vr = VideoReader(item["video_path"])
frames = vr.get_batch([0, 8, 16])
return frames, item["labels"]
loader = DataLoader(BreakfastDataset(ds), batch_size=4)๐ข Splits Description
The dataset is partitioned by participant ID:
Each split has its own metadata JSONL file.
๐ Citation
If you use the Breakfast Actions dataset, please cite:
@inproceedings{kuehne2014language,
title={The language of actions: Recovering the syntax and semantics of goal-directed human activities},
author={Kuehne, Hildegard and Arslan, Ali and Serre, Thomas},
booktitle={Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition},
pages={780--787},
year={2014}
}