Team Ai
Datasetpublic

thxplz/HowToEat-test

HowToEat: Hand-Object Interaction and Eating Action in Eating Scenarios HowToEat is an image dataset for analysing eating behaviour. It provides: Hand-object interaction + eating face detection (hand_object_detection): 95,190 images with 190,333 hand instances (box, left/right side, contact state, and the box and category of the held object) and 151,620 face instances (box, eating / not eating). Eating action recognition (eating_recognition): 6,280 manually labelled faces… See the full description on the dataset page: https://huggingface.co/datasets/thxplz/HowToEat-test.

sourceHugging Faceotherupdated 5h agoView on Hugging Face
0likes45downloads
Dataset Card

HowToEat: Hand-Object Interaction and Eating Action in Eating Scenarios

HowToEat is an image dataset for analysing eating behaviour. It provides:

  1. 1.Hand-object interaction + eating face detection (hand_object_detection): 95,190 images with 190,333 hand instances (box, left/right side, contact state, and the box and category of the held object) and 151,620 face instances (box, eating / not eating).
  2. 2.Eating action recognition (eating_recognition): 6,280 manually labelled faces (eating / not eating) for image classification.

The images are frames from 6,701 publicly available eating videos covering 12 eating and drinking scenarios.

  • —Paper: Yingcheng Wang, Junwen Chen, Keiji Yanai. HowToEat: Exploring Human-Object Interaction and Eating Action in Eating Scenarios. MADiMa '23 (8th International Workshop on Multimedia Assisted Dietary Management, in conjunction with ACM Multimedia 2023). doi:10.1145/3607828.3617790
  • —Institution: Department of Informatics, The University of Electro-Communications, Tokyo, Japan
  • —License: HowToEat Research-Only License: non-commercial research and education only (see License)
  • —Not included: the source videos and the trained models are not released.

Quick start

Access is gated: accept the license on this page, then log in with huggingface-cli login.

python
from datasets import load_dataset

# Task 1: hand-object interaction + eating face detection
det = load_dataset("thxplz/HowToEat-test", "hand_object_detection")
sample = det["train"][0]
sample["image"]          # PIL.Image, 1920x1080
sample["hands"]          # list of hand instances
sample["faces"]          # list of face instances

# Task 2: eating action recognition
rec = load_dataset("thxplz/HowToEat-test", "eating_recognition")
rec["train"][0]["face_crop"], rec["train"][0]["label"]   # 224x224 crop, 0 = eating

# Stream instead of downloading everything (~35 GB)
det_stream = load_dataset("thxplz/HowToEat-test", "hand_object_detection", split="test", streaming=True)

Dataset structure

HowToEat/
├── hand_object_detection/       # Parquet shards (images embedded), Task 1
├── eating_recognition/          # Parquet shards (images embedded), Task 2
├── annotations/                 # the same annotations in the original JSON format
│   ├── detection_train.json     #   paper split
│   ├── detection_test.json
│   ├── detection_train_balanced.json   # balanced split (see "Splits")
│   ├── detection_test_balanced.json
│   ├── eating_recognition_train.json
│   └── eating_recognition_test.json
├── metadata/
│   └── categories.json          # id <-> name maps
├── LICENSE.md
└── README.md

Image files are identified by file_name = "<verb>/<category>/<video_id>_<frame>.jpg", for example eating/pizza/v3174_011858.jpg. <video_id> is an anonymous video ID (v0001 … v6701) and <frame> is the 6-digit frame index within that video. Frames with the same video_id come from the same video.

Scene categories (12)

hamburger, beer (under drinking/), bread, pizza, pasta, noodles, with_knife_and_fork, sushi, with_spoon, with_fork, sandwich, with_chopsticks. All except beer are under eating/. The category describes the scene of the whole video; it is not a per-instance label.

Config hand_object_detection

ColumnTypeDescription
imageImageFull frame, 1920×1080 JPEG
file_namestringKey shared with the JSON annotations
video_idstringAnonymous video ID (e.g. v0123)
frameintFrame index in the source video
categoryClassLabel (12)Scene category of the video
width, heightintImage size
handslist of structOne entry per hand, see below
faceslist of structOne entry per face, see below
split_balancedstring"train" / "test" in the balanced split (see Splits)

hands[i]:

FieldDescription
hand_box[x_min, y_min, x_max, y_max], absolute pixels
hand_side0 = left, 1 = right
ho_exist1 = the hand is in contact with a portable object, 0 = no contact
obj_boxBox of the held object, [] when ho_exist = 0
obj_catObject category (table below), -1 when ho_exist = 0

faces[j]:

FieldDescription
face_box[x_min, y_min, x_max, y_max], absolute pixels
eat1 = eating, 0 = not eating

Object categories (obj_cat):

idnameinstancesidnameinstances
0food58,7207bottle6,282
1chopsticks34,6638cup10,594
2fork9,2689glass9,742
3spoon10,78310can875
4knife12,09811napkin7,598
5bowl3,65612unknown19
6plate1,121−1(no contact)24,914

Example record (original JSON format in annotations/detection_*.json):

json
{
  "file_name": "eating/pizza/v3174_011858.jpg",
  "hand_obj": [
    {"hand_box": [928, 406, 1062, 519], "hand_side": 0, "ho_exist": 1,
     "obj_box": [931, 391, 1005, 453], "obj_cat": 0},
    {"hand_box": [900, 417, 979, 550], "hand_side": 1, "ho_exist": 1,
     "obj_box": [922, 390, 1006, 478], "obj_cat": 0}
  ],
  "face": [{"face_box": [903, 264, 1057, 411], "eat": 1}]
}

(In the JSON files, the hand list is called hand_obj and the face list face.)

Config eating_recognition

ColumnTypeDescription
imageImageFull frame the face was taken from
face_cropImage224×224 square crop around face_box with 15% context (the crop used at test time in the paper)
file_name, video_id, frame, categoryAs above
face_boxlist[int][x_min, y_min, x_max, y_max] of the labelled face
labelClassLabel`0` = eating, `1` = not_eating
⚠️ Label polarity differs between the two configs. In eating_recognition, label = 0 means eating. In hand_object_detection, eat = 1 means eating. We kept the conventions of the original training code for both.

Splits

hand_object_detection

The train / test splits of this config are the split used in the paper (baseline results below). No video appears in both splits. Because the split was made in category order, the test set does not cover all categories: hamburger, beer, bread, pizza, pasta, noodles and sandwich appear only in train.

For a test set that covers every category, we also provide a balanced split: a random 4:1 split by video, where each category has about 11–24% of its images in test. It is available as the split_balanced column and as annotations/detection_*_balanced.json. No baseline results have been published on the balanced split.

python
from datasets import load_dataset, concatenate_datasets
det = load_dataset("thxplz/HowToEat-test", "hand_object_detection")
full = concatenate_datasets([det["train"], det["test"]])
train_bal = full.filter(lambda s: s == "train", input_columns="split_balanced")
test_bal  = full.filter(lambda s: s == "test",  input_columns="split_balanced")
Paper split (train / test)Balanced split (train / test)
Images76,905 / 18,28576,025 / 19,165
Videos5,550 / 1,1045,351 / 1,303
Hands155,634 / 34,699153,532 / 36,801
Faces122,183 / 29,437122,033 / 29,587

Images per category:

CategoryTotalPaper trainPaper testBalanced trainBalanced test
hamburger19,23919,239015,5203,719
sushi15,8432,79913,04412,5113,332
pizza14,22114,221011,8202,401
beer11,70111,70109,2582,443
noodles10,89110,89108,5112,380
bread7,6447,64406,0811,563
pasta7,4657,46505,6641,801
sandwich2,8342,83402,258576
with_chopsticks1,727351,6921,359368
with_spoon1,58221,5801,290292
withknifeand_fork1,165711,094974191
with_fork878387577999
Total95,19076,90518,28576,02519,165

eating_recognition

SplitFacesEatingNot eatingVideos
train5,0333,1081,925537
test1,247770477353

The split is stratified 4:1 within each (category, label) group, as in the paper. It was made per image, not per video, so 332 videos have frames in both train and test. Keep this in mind when you interpret test accuracy.


Statistics (hand_object_detection, both splits combined)

  • —Hands: 190,333 (left 85,797 / right 104,536). In contact with an object: 165,419. No contact: 24,914.
  • —Faces: 151,620 (eating 69,251 / not eating 82,369). Every image has 1–5 faces. A few images contain faces but no annotated hands.
  • —Faces per image: 1 face 56,153 images · 2 faces 27,518 · 3 faces 7,142 · 4 faces 2,880 · 5 faces 1,497.

Baseline results

Results reported in the paper on the paper split (test set), with mAP at IoU > 0.5 (see the paper §5.1 for the hand-object matching rule).

SOV-STG-H2E-S (multi-task, single model):

Left hand: no contactLeft hand: portable objectRight hand: no contactRight hand: portable object**Hand-object mAP**Face: not eatingFace: eating**Face mAP**
61.9187.7947.9888.5671.5657.4373.8965.66

Eating action recognition (eating_recognition), ResNet-50 (ImageNet-1K pre-trained, fine-tuned): 86.4% test accuracy.

The trained models are not released.


How the dataset was built

  1. 1.Video collection. Eating and drinking videos were collected for 12 scenarios (e.g. eating hamburger, drinking beer, eating with a spoon). The videos themselves are not distributed.
  2. 2.Frame extraction. A PPDM hand-object interaction detector trained on 100DOH and a RetinaFace (ResNet-50) face detector were run at 1 frame per second. Frames where a hand-held object overlapped the mouth landmarks were kept: 99,903 frames.
  3. 3.Eating labels for faces. 6,280 face crops were labelled manually (eating / not eating). This is the eating_recognition config. A ResNet-50 classifier trained on them then labelled all faces automatically.
  4. 4.Hand-object annotation. An SOV-STG-Hand model (trained on 100DOH, re-categorised into no contact / portable object) produced hand, side, contact and object boxes. Object categories were added and all annotations were then checked and corrected manually with the VIA annotation tool. Frames that could not be annotated reliably were marked invalid.
  5. 5.Filtering. We removed images without faces, images with more than 5 faces, and images whose largest face is smaller than 400 px. Faces smaller than 300 px were removed. Together with the invalid frames, this reduced the 97,484 verified frames to 95,190 images.

Labelling rules for eating faces: a face is eating if the person is performing the act of eating (mouth open with food or utensil entering it, or the face clearly shows eating while an object covers the mouth). An object in front of a closed mouth counts as not eating. Images that cannot be judged were not labelled.


Limitations and biases

  • —Partly automatic labels. Face boxes come from RetinaFace. Eating labels on detection faces come from a classifier, and hand-object boxes were pre-annotated by a model before manual checking. Some errors remain (see the paper, Fig. 6c).
  • —Domain. The videos are mostly eating vlogs filmed for an audience: frontal faces, often a single person, good lighting. The foods are limited to the 12 scenario categories, and the people and regions shown are not representative of the world population.
  • —Boxes outside the image. Some boxes extend beyond the image border (for example, negative coordinates for a face cut off at the top), mostly face boxes produced by the face detector. They are kept exactly as used in the paper; clip them to [0, width] × [0, height] if your code requires it.
  • —Category imbalance. unknown (id 12) has only 19 instances and can 875. The with_* categories are much smaller than the food categories.
  • —Paper split coverage. See Splits: the paper's test set covers only 5 of the 12 categories in meaningful numbers.

Ethical considerations

The images show real, identifiable people taken from publicly available videos. The dataset is intended only for research on eating behaviour and dietary assessment. It must not be used for face recognition, identification, tracking or profiling of individuals (see the license).

Removal requests. If you appear in the dataset or own one of the source videos and want content removed, contact `<contact email>` with the file_name or video_id. We will remove it from this repository.


License

The annotations, metadata and scripts are released under the [HowToEat Research-Only License](LICENSE.md). In short:

  • —✅ Non-commercial research and education
  • —❌ Commercial use of any kind, including training commercial models
  • —❌ Redistribution or re-hosting (a few example images in publications are fine)
  • —❌ Face recognition, identification or surveillance
  • —📌 Citation of the paper is required

The images are frames from publicly available online videos. Their copyright belongs to the original video owners, and the authors do not grant any rights to them. They are made available for research use only, to the extent permitted by applicable law.


Citation

bibtex
@inproceedings{wang2023howtoeat,
  title     = {{HowToEat}: Exploring Human Object Interaction and Eating Action in Eating Scenarios},
  author    = {Wang, Yingcheng and Chen, Junwen and Yanai, Keiji},
  booktitle = {Proceedings of the 8th International Workshop on Multimedia Assisted Dietary Management},
  series    = {MADiMa '23},
  pages     = {71--78},
  year      = {2023},
  publisher = {Association for Computing Machinery},
  address   = {New York, NY, USA},
  location  = {Ottawa, ON, Canada},
  doi       = {10.1145/3607828.3617790},
  url       = {https://doi.org/10.1145/3607828.3617790}
}

Acknowledgments

This work was supported by JSPS KAKENHI Grant Numbers 21H05812, 22H00540, 22H00548, and 22K19808.