Team Ai
Datasetpublic

CVML-TueAI/Breakfast-Actions

๐Ÿณ Breakfast Actions Dataset (HF + WebDataset Ready) This repository hosts the Breakfast Actions dataset metadata and videos, organized for modern deep learning workflows.It provides: 4 evaluation splits (s1, s2, s3, s4) JSONL metadata describing each video, participant, camera, and frame-level action segments Raw AVI videos stored directly on HuggingFace Optional WebDataset shards for streaming training ๐Ÿ“ Folder Layout Breakfast-Actions/ โ”‚ โ”œโ”€โ”€โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/CVML-TueAI/Breakfast-Actions.

sourceHugging Faceupdated 10mo agoView on Hugging Face
0likes1.7kdownloads
Dataset Card

๐Ÿณ Breakfast Actions Dataset (HF + WebDataset Ready)

This repository hosts the Breakfast Actions dataset metadata and videos, organized for modern deep learning workflows. It provides:

  • โ€”4 evaluation splits (s1, s2, s3, s4)
  • โ€”JSONL metadata describing each video, participant, camera, and frame-level action segments
  • โ€”Raw AVI videos stored directly on HuggingFace
  • โ€”Optional WebDataset shards for streaming training

๐Ÿ“ Folder Layout

Breakfast-Actions/
โ”‚
โ”œโ”€โ”€ Converted_Data/
โ”‚     โ”œโ”€โ”€ metadata_s1.jsonl
โ”‚     โ”œโ”€โ”€ metadata_s2.jsonl
โ”‚     โ”œโ”€โ”€ metadata_s3.jsonl
โ”‚     โ””โ”€โ”€ metadata_s4.jsonl
โ”‚
โ”œโ”€โ”€ Videos/
โ”‚     โ”œโ”€โ”€ P03/cam01/*.avi
โ”‚     โ”œโ”€โ”€ P03/cam02/*.avi
โ”‚     โ”œโ”€โ”€ P04/cam01/*.avi
โ”‚     โ””โ”€โ”€ ... (participants P03โ€“P54, multiple cameras)
โ”‚
โ””โ”€โ”€ WebDataset_Shards/   (optional)
       โ”œโ”€โ”€ 000000.tar
       โ”œโ”€โ”€ 000001.tar
       โ””โ”€โ”€ ...

๐Ÿ“ JSONL Record Format

Each metadata line looks like:

json
{
  "video_path": "Videos/P03/cam01/P03_coffee.avi",
  "participant": "P03",
  "camera": "cam01",
  "video": "P03_coffee",
  "labels": [
      {"start": 1, "end": 385, "label": "SIL"},
      {"start": 385, "end": 599, "label": "pour_oil"},
      ...
  ]
}

All video paths match the directory structure inside the HF repo.


๐Ÿ”น Load Metadata Using HuggingFace Datasets

python
from datasets import load_dataset

ds = load_dataset("json", data_files="metadata_s2.jsonl")["train"]

# Select all videos belonging to split s2
subset = ds

๐Ÿ”น Load and Decode a Video

Using Decord

python
from decord import VideoReader
item = ds[0]

vr = VideoReader(item["video_path"])
frame0 = vr[0]   # first frame

Using TorchVision

python
from torchvision.io import read_video
video, audio, info = read_video(item["video_path"])

๐Ÿ”น WebDataset Version (Optional)

If the dataset includes .tar shards:

python
import webdataset as wds, jsonlines

ids = [rec["video_path"] for rec in jsonlines.open("metadata_s2.jsonl")]
dset = wds.WebDataset("WebDataset_Shards/*.tar").select(lambda s: s["json"]["video_path"] in ids)

Each shard contains:

  • โ€”xxx.avi โ†’ video bytes
  • โ€”xxx.json โ†’ metadata JSON

๐Ÿ”น PyTorch Example

python
from torch.utils.data import Dataset, DataLoader
from decord import VideoReader

class BreakfastDataset(Dataset):
    def __init__(self, subset): self.subset = subset
    def __len__(self): return len(self.subset)
    def __getitem__(self, idx):
        item = self.subset[idx]
        vr = VideoReader(item["video_path"])
        frames = vr.get_batch([0, 8, 16])
        return frames, item["labels"]

loader = DataLoader(BreakfastDataset(ds), batch_size=4)

๐Ÿ”ข Splits Description

The dataset is partitioned by participant ID:

SplitParticipants
s1P03โ€“P15
s2P16โ€“P28
s3P29โ€“P41
s4P42โ€“P54

Each split has its own metadata JSONL file.


๐Ÿ“š Citation

If you use the Breakfast Actions dataset, please cite:

bibtex
@inproceedings{kuehne2014language,
  title={The language of actions: Recovering the syntax and semantics of goal-directed human activities},
  author={Kuehne, Hildegard and Arslan, Ali and Serre, Thomas},
  booktitle={Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition},
  pages={780--787},
  year={2014}
}