Sonam5/Calf-Play-Behavior-Dataset
Video recordings and play behaviour annotations of group-housed dairy calves on two Dutch farms Status: version 1.0. A Data Descriptor describing this dataset is in preparation for Scientific Data; the citation below will be updated when it is published. Summary This dataset contains video of group-housed dairy calves recorded during four camera deployments of about one week each on two farms in the Netherlands, together with continuous per-calf behaviour… See the full description on the dataset page: https://huggingface.co/datasets/Sonam5/Calf-Play-Behavior-Dataset.
Video recordings and play behaviour annotations of group-housed dairy calves on two Dutch farms
Status: version 1.0. A Data Descriptor describing this dataset is in preparation for Scientific Data; the citation below will be updated when it is published.
Summary
This dataset contains video of group-housed dairy calves recorded during four camera deployments of about one week each on two farms in the Netherlands, together with continuous per-calf behaviour annotations focused on play.
Repository structure
README.md this card
LICENSE CC0 1.0 Universal legal code
metadata/
sessions.csv one row per annotation or frame session (9)
videos.csv one row per video file (1,175): times, size, codec, resolution, fps, notes
calves.csv one row per calf (12) with every spelling of its name
ethogram.csv raw label -> ethogram behaviour -> three-class label; definitions; event counts
known_issues.md every verified anomaly (read this first)
files_manifest.csv path, bytes, sha256 of every file
frame_index/<session_id>.csv.gz OCR camera-clock time of every decoded frame (7 files) + README.md (OCR method)
qc/ quality-control tables of the frame index and the labels + README.md
videos/<visit>/<YYYY-MM-DD>/<original file name>.mp4 visit = Tol | Tol2 | Tol3 | Eem (1,175 files)
frames/<session_id>/<session_id>_video<NNN>.tar one tar per 3,000-frame folder (372 files)
annotations/
README.md what the two annotation layers are, how they were made, all corrections
event_logs/<session_id>.csv behaviour event logs, one per session (9): the primary annotation
labels/<session_id>.csv.gz one row per labelled frame x calf (4 sessions)
tracking/<session_id>.tar <session_id>/masks/video<NNN>/*.json + <session_id>/initial_prompts/*.txt (7 files)Total size about 2.37 TB: videos 1,782.0 GB, frame shards 575.8 GB, tracking 8.0 GB, everything else about 7 MB.
Sessions
Times are local camera time (CEST, UTC+2), as shown on the video overlay.
Time alignment: use the camera clock, not the frame index
The camera recording drops frames (evidence in metadata/known_issues.md, item 9), so a frame's wall-clock time cannot be computed from its index, its position in the file or the nominal frame rate. Each decoded frame therefore has the time shown on the camera clock burned into the image, read by Tesseract OCR with the original pipeline (metadata/frame_index/):
Successvalues are not corrected. Between 0% and 3.8% of them per session are misreads (for example year 2624, or a date two days off; counts inmetadata/qc/ocr_plausibility.csv). Screen them against neighbouring frames, e.g. drop values more than 60 s from a rolling median.- Coverage is low for the 2024 Tol2 sessions (overlay unreadable in bright daylight):
Tol2_2024-06-11_AMhas aSuccessreading in only 12.6% of its minutes, andTol2_2024-06-11_PMnone before 17:45. Annotations in those periods cannot be aligned to frames by the clock. - Several frames share one clock second, so clock-aligned labels have one-second resolution.
Annotations
There are two annotation layers (details, column definitions and every correction in annotations/README.md):
Event logs. One Observer XT observation per video file (about 34 min). For every calf (Subject, with the dataset-wide calf_id), the observer coded the current state with State start/State stop events; State point rows are instantaneous markers. The camera-clock time of an event is observation_start + Time_Relative_sf (seconds). Do not use Date_Time_Absolute_dmy_hmsf, Date_dmy or Time_Absolute_*: they hold the observer's computer clock at the time of coding. The observers' exports were checked and corrected by the authors before release (placeholder rows, a calf-name variant, three mistyped observation dates, merging of partial exports); no behaviour, time or duration was changed, and annotations/README.md lists every change.
Ethogram (metadata/ethogram.csv):
Head shakes were scored as play only directly before or after another play behaviour, kicks only in a play or social context; a kick was scored together with a buck when both occurred simultaneously (buck-kick). Definitions are in ethogram.csv.
Labels (annotations/labels/, 57,920 rows): one row per frame and calf for Tol_2024-04-24 (TOL-2642, TOL-2643, TOL-2644), Tol2_2024-06-11_AM and Tol2_2024-06-11_PM (TOL-2642, TOL-2644) and Eem_2024-06-19 (EEM-9, 06:03–11:23).
The labels cover only these calves, frames with a Success OCR reading, and the parts of the day kept by the labelling pipeline (long rest periods between play bouts were not carried over, so most Not Playing time is unlabelled). Frames in which the calf is out of view are excluded, and Buck, Kick, Jump and Turn never decide a label. For other calves and frames, derive labels from the event logs and the frame index (example below). Time inside an annotated window with no coded state, and all time outside annotated windows, is unlabelled, not "Not Playing".
Tracking (tracking/<session_id>.tar): per calf and 3,000-frame folder, a JSON file {frame name: {"bounding_box": [x, y, w, h], "contour": [...]}}, and the initial prompt boxes (calf name: x, y, w, h). Tracked: Tol_2024-04-24 3 calves; Tol2_2024-06-11_AM/_PM 2 calves; Eem_2024-06-18 5 calves (folders 19–61); Eem_2024-06-19 calf 9. Tracks are automatic SAMURAI output and were not manually corrected; identity switches were not audited.
How to load
REPO is the repository ID (Sonam5/Calf-Play-Behavior-Dataset).
Download metadata and annotations only (a few MB):
from huggingface_hub import snapshot_download
REPO = "Sonam5/Calf-Play-Behavior-Dataset"
local = snapshot_download(REPO, repo_type="dataset",
allow_patterns=["README.md", "LICENSE", "metadata/*", "annotations/*"])Download the frames of one session (e.g. 92 GB for Tol2_2024-06-11_PM):
snapshot_download(REPO, repo_type="dataset", allow_patterns=["frames/Tol2_2024-06-11_PM/*"], local_dir="calves")Read one frame from a shard:
import io, tarfile
from PIL import Image
from huggingface_hub import hf_hub_download
path = hf_hub_download(REPO, "frames/Tol_2024-04-24/Tol_2024-04-24_video001.tar", repo_type="dataset")
with tarfile.open(path) as tf:
img = Image.open(io.BytesIO(tf.extractfile("0000001.jpg").read()))Stream shards with the webdataset library (each sample has the key 0000001 and the field jpg):
import webdataset as wds
urls = f"https://huggingface.co/datasets/{REPO}/resolve/main/frames/Tol_2024-04-24/Tol_2024-04-24_video{{001..030}}.tar"
for sample in wds.WebDataset(urls, shardshuffle=False).decode("pil"):
key, img = sample["__key__"], sample["jpg"]Join the labels with the frame index:
import pandas as pd
sid = "Tol_2024-04-24"
lab = pd.read_csv(f"{local}/annotations/labels/{sid}.csv.gz")
idx = pd.read_csv(f"{local}/metadata/frame_index/{sid}.csv.gz")
lab = lab.merge(idx[["frame_name", "chunk", "ocr_status"]], on="frame_name")Label frames from the event logs (any calf, any aligned frame; example for one session):
sid = "Tol2_2024-06-12"
ev = pd.read_csv(f"{local}/annotations/event_logs/{sid}.csv")
st = ev[ev["Event_Type"] == "State start"].copy()
st["start"] = pd.to_datetime(st["observation_start"]) + pd.to_timedelta(st["Time_Relative_sf"], unit="s")
st["end"] = st["start"] + pd.to_timedelta(st["Duration_sf"], unit="s")
idx = pd.read_csv(f"{local}/metadata/frame_index/{sid}.csv.gz")
idx = idx[idx["ocr_status"] == "Success"].copy()
idx["t"] = pd.to_datetime(idx["ocr_timestamp"], errors="coerce") + pd.Timedelta(seconds=0.5) # middle of the clock second
# screen OCR misreads before use (e.g. rolling-median check), then, per calf:
calf = st[st["calf_id"] == "TOL-2644"]
for _, e in calf.iterrows(): # states can overlap: Buck, Kick, Jump, Turn, Milk Feeding and Management are coded on top of the main state
idx.loc[(idx["t"] >= e["start"]) & (idx["t"] < e["end"]), e["Behavior"]] = Trueobservation_start already contains the corrected start of every observation (for example the 24 observations of Eem_2024-06-20 whose typed start is missing or mistyped). Tol_2024-04-24 start times are 2 s earlier than the file names; see metadata/known_issues.md.
Usage notes
- Consecutive frames are nearly identical and several frames share one clock second. When training or evaluating models, split by blocks of time, by session or by farm rather than by random frames.
- The April and June 2024 Tol sessions share calves, and the two halves of 11 June 2024 are the same calves on the same day: treat all Tol sessions as one farm when comparing farms.
- The released frames, labels and tracks cover the annotated sessions and calves listed above. The videos of all 1,175 files are released, so users can decode further frames and run the OCR and tracking code (see Code availability) on other periods or calves.
Known issues (summary)
Read metadata/known_issues.md before use. The most important are:
- The April 2024 recording covers only about 05:40–24:00 daily; no video on 20–21 April 2024; no video 13:12:54–13:34:27 on 11 June 2024.
- Two videos cannot be played (no MP4 index):
Tol_ch01_20240418-113301-121506.mp4(corrupt) andTol2_ch04_20240611_164646_172030.mp4(truncated). They are released and flagged. - Tol3 files spanning midnight are named by their end date and each contains a recording gap of about 30 min.
Tol2_2024-06-11_AMframes 1–15,000 (06:09–06:59) were deleted;Eem_2024-06-18has no frames for about 63 annotated minutes; the pilot day andEem_2024-06-20were not decoded. Two single frames are missing.- OCR coverage is low for the 2024 Tol sessions and
Successvalues include misreads (see above); 461 label rows (90 of the 452 rows ofTol2_2024-06-11_AM) sit on implausible readings. - Event times use the position in the video file; because frames are dropped, they can differ from the camera clock by up to a few seconds towards the end of a file.
- Within the annotated windows, 2.8% of calf-time in
Tol_2024-04-24and 1.5% inEem_2024-06-19has no coded state (unlabelled). - One observer per observation; inter-observer reliability was not measured.
- Channel labels differ between file names and overlays (Eem files
ch04, overlayCH3; Tol2 annotation files sayCh01, camera is channel 4). - Calves in neighbouring pens are visible in the Tol2 and Tol3 views; only the pen's own calves are annotated.
Not included
Recordings of the other farms of the observational study (no open-release agreement); leg-accelerometer data; per-calf image crops; decoded frames of two unannotated days; intermediate processing tables, which the event logs, labels and frame index supersede.
Ethics and consent
- The observational study was approved by the Animal Welfare Body of Utrecht University (approval 10818-2024-02). Recording was non-invasive.
- The farm owners consented to recording and to public release of these data under CC0 (no rights reserved).
- Farm staff occasionally appear in the footage; faces are not recognisable.
- Observer initials were replaced by
O1–O5in all annotation files. Farms are identified by the short codesTol(withTol2andTol3denoting later recording rounds at the same farm) andEem; no personal data of farmers are included. - If you find identifiable personal information in the data, please contact us (below) so that it can be removed in a new revision.
Code availability
- https://github.com/Sonam525/Play-Behavior-Dataset: the code used to build this release (frame decoding, camera-clock OCR, tracking, labels and quality-control tables).
- https://github.com/Sonam525/Individual-Behavior-Analysis-with-CV: the video-analysis pipeline of the related play-behaviour study.
Licence
CC0 1.0 Universal. You may copy, modify and use the data for any purpose without asking permission. We ask, but do not require, that you cite the Data Descriptor and this dataset.
Citation
Until the Data Descriptor is published, please cite this dataset:
@misc{calfplay_dataset_2026,
author = {Yang, Haiyu and Lesscher, Heidi and Liu, Enhong and Hostens, Miel},
title = {Video recordings and play behaviour annotations of group-housed dairy calves on two Dutch farms},
year = {2026},
publisher = {Hugging Face},
version = {1.0},
doi = {10.57967/hf/10695},
url = {https://huggingface.co/datasets/Sonam5/Calf-Play-Behavior-Dataset}
}Related preprints by the authors: Yang et al., arXiv:2602.00111 (2026); Yang et al., arXiv:2509.12047 (2025); Yang et al., arXiv:2603.17782 (2026).
Contact and maintenance
Corresponding author: Haiyu Yang, Department of Animal Science, Cornell University, Ithaca, NY, USA. Please open a discussion in the repository's Community tab.
Changelog
- Version 1.0 (2026-09-30): first complete release. The revision cited by the Data Descriptor is fixed under its DOI; later revisions get a new DOI and an entry here.
