Team Ai
Datasetpublic

Sonam5/Calf-Play-Behavior-Dataset

Video recordings and play behaviour annotations of group-housed dairy calves on two Dutch farms Status: version 1.0. A Data Descriptor describing this dataset is in preparation for Scientific Data; the citation below will be updated when it is published. Summary This dataset contains video of group-housed dairy calves recorded during four camera deployments of about one week each on two farms in the Netherlands, together with continuous per-calf behaviour… See the full description on the dataset page: https://huggingface.co/datasets/Sonam5/Calf-Play-Behavior-Dataset.

sourceHugging Facecc0-1.0updated 9d agoView on Hugging Face
0likes245downloads
Dataset Card

Video recordings and play behaviour annotations of group-housed dairy calves on two Dutch farms

Status: version 1.0. A Data Descriptor describing this dataset is in preparation for Scientific Data; the citation below will be updated when it is published.

Summary

This dataset contains video of group-housed dairy calves recorded during four camera deployments of about one week each on two farms in the Netherlands, together with continuous per-calf behaviour annotations focused on play.

FarmsTol: a research and teaching dairy farm in the Netherlands (3 deployments: April 2024 pilot, June 2024, May 2025). Eem: a commercial dairy farm (June 2024).
Calves12 (Tol: 4 in April and June 2024, three of them the same calves; 2 in May 2025. Eem: 5)
Video1,175 MP4 files, 652.6 h, 1,782.0 GB (one camera per deployment; 1280×720 to 2304×1296 pixels; 25 or 12 frames/s)
AnnotationsBehaviour event logs (Noldus The Observer XT) for 8 observation days: 89.96 h of video with every calf coded continuously (event times to the video frame, 18-item ethogram), plus a 13.78-h pilot day with one focal calf
Frames1,106,110 JPEG frames (every 5th video frame) from 7 annotated sessions, in 372 tar shards (575.8 GB)
Frame timeper-frame camera-clock time read by OCR from the burned-in overlay (metadata/frame_index/)
DerivedSAMURAI (SAM 2.1) boxes and masks for tracked calves; frame-level three-class play labels for 57,920 calf-frames
Sizeabout 2.37 TB in total
LicenceCC0 1.0 (public domain dedication)

Repository structure

README.md                         this card
LICENSE                           CC0 1.0 Universal legal code
metadata/
  sessions.csv                    one row per annotation or frame session (9)
  videos.csv                      one row per video file (1,175): times, size, codec, resolution, fps, notes
  calves.csv                      one row per calf (12) with every spelling of its name
  ethogram.csv                    raw label -> ethogram behaviour -> three-class label; definitions; event counts
  known_issues.md                 every verified anomaly (read this first)
  files_manifest.csv              path, bytes, sha256 of every file
  frame_index/<session_id>.csv.gz OCR camera-clock time of every decoded frame (7 files) + README.md (OCR method)
  qc/                             quality-control tables of the frame index and the labels + README.md
videos/<visit>/<YYYY-MM-DD>/<original file name>.mp4      visit = Tol | Tol2 | Tol3 | Eem  (1,175 files)
frames/<session_id>/<session_id>_video<NNN>.tar           one tar per 3,000-frame folder (372 files)
annotations/
  README.md                       what the two annotation layers are, how they were made, all corrections
  event_logs/<session_id>.csv     behaviour event logs, one per session (9): the primary annotation
  labels/<session_id>.csv.gz      one row per labelled frame x calf (4 sessions)
tracking/<session_id>.tar         <session_id>/masks/video<NNN>/*.json + <session_id>/initial_prompts/*.txt (7 files)

Total size about 2.37 TB: videos 1,782.0 GB, frame shards 575.8 GB, tracking 8.0 GB, everything else about 7 MB.

Sessions

Times are local camera time (CEST, UTC+2), as shown on the video overlay.

Session IDAnnotated window(s)Calves codedAnnotated hDecoded framesShardsOCR `Success` frames
Tol_2024-04-18 (pilot)09:27–12:15, 12:57–23:551 focal13.78–––
Tol_2024-04-2405:48–10:4644.9789,5713034.7%
Tol2_2024-06-11_AM06:09–13:1247.05111,860383.0%
Tol2_2024-06-11_PM13:34–23:59 (not 16:46–17:20)49.84177,1116046.7%
Tol2_2024-06-1206:03–10:3344.4980,8842735.7%
Tol3_2025-05-1206:21–23:01216.65143,8804899.6%
Eem_2024-06-1812:51–13:00, 13:03–23:59511.07180,4856199.8%
Eem_2024-06-1906:03–23:59 (not 22:02–22:30)517.43322,31910899.8%
Eem_2024-06-2006:01–00:30 (+1 day)518.45–––

Time alignment: use the camera clock, not the frame index

The camera recording drops frames (evidence in metadata/known_issues.md, item 9), so a frame's wall-clock time cannot be computed from its index, its position in the file or the nominal frame rate. Each decoded frame therefore has the time shown on the camera clock burned into the image, read by Tesseract OCR with the original pipeline (metadata/frame_index/):

ColumnMeaning
session_idsession ID (table above)
frame_nameJPEG name = 7-digit frame index within the session (0015001.jpg)
chunkoriginal 3,000-frame folder (video6 -> shard <session_id>_video006.tar)
ocr_timestampYYYY-MM-DDTHH:MM:SS, local camera time; empty unless ocr_status is Success
ocr_statusSuccess (the text parsed as a date-time) or ParseError
sourceoriginal OCR table, or re-run of the same code (sessions whose original table was missing or overwritten)
  • —Success values are not corrected. Between 0% and 3.8% of them per session are misreads (for example year 2624, or a date two days off; counts in metadata/qc/ocr_plausibility.csv). Screen them against neighbouring frames, e.g. drop values more than 60 s from a rolling median.
  • —Coverage is low for the 2024 Tol2 sessions (overlay unreadable in bright daylight): Tol2_2024-06-11_AM has a Success reading in only 12.6% of its minutes, and Tol2_2024-06-11_PM none before 17:45. Annotations in those periods cannot be aligned to frames by the clock.
  • —Several frames share one clock second, so clock-aligned labels have one-second resolution.

Annotations

There are two annotation layers (details, column definitions and every correction in annotations/README.md):

LayerUse it for
annotations/event_logs/The primary annotation. Every behaviour the observers coded for every calf in the pen, with start time and duration, for all 9 sessions. Use it for time budgets, bout analyses, any calf or session, and to derive your own frame labels.
annotations/labels/Ready-made frame-level three-class play labels (Active Playing / Non Active Playing / Not Playing) for the calves processed by the labelling pipeline in 4 sessions.

Event logs. One Observer XT observation per video file (about 34 min). For every calf (Subject, with the dataset-wide calf_id), the observer coded the current state with State start/State stop events; State point rows are instantaneous markers. The camera-clock time of an event is observation_start + Time_Relative_sf (seconds). Do not use Date_Time_Absolute_dmy_hmsf, Date_dmy or Time_Absolute_*: they hold the observer's computer clock at the time of coding. The observers' exports were checked and corrected by the authors before release (placeholder rows, a calf-name variant, three mistyped observation dates, merging of partial exports); no behaviour, time or duration was changed, and annotations/README.md lists every change.

Ethogram (metadata/ethogram.csv):

ClassRaw labels
Active Playinggallop, leap, leap sideways, Buck, Kick/kick, Jump/jump, Turn/turn
Non Active Playingfrontal pushing, mount, head-shake, head-butting (object), butting fixtures, butting/rubbing straw, bouncing / play-bouncing (not in the ethogram)
Not Playingrest behavior (background state for any non-play behaviour), Milk Feeding / voer krijgen, Management (someone in the pen / stalls)
excluded from the labelsout of view
not a calf labelIndividuall (point marker: a person or animal in front of the pen, copied to every calf); empty point markers

Head shakes were scored as play only directly before or after another play behaviour, kicks only in a play or social context; a kick was scored together with a buck when both occurred simultaneously (buck-kick). Definitions are in ethogram.csv.

Labels (annotations/labels/, 57,920 rows): one row per frame and calf for Tol_2024-04-24 (TOL-2642, TOL-2643, TOL-2644), Tol2_2024-06-11_AM and Tol2_2024-06-11_PM (TOL-2642, TOL-2644) and Eem_2024-06-19 (EEM-9, 06:03–11:23).

ColumnMeaning
session_id, frame_nameframe
frame_tarshard containing the frame
calf_idcalf (metadata/calves.csv)
subjectcalf name as in the event logs
ocr_timestampOCR clock second of the frame (as in metadata/frame_index/)
primary_behaviour, secondary_behaviourmost and second most frequent main state of this calf in this second
labelActive Playing if either is Active; else Non Active Playing if either is Non-Active; else Not Playing

The labels cover only these calves, frames with a Success OCR reading, and the parts of the day kept by the labelling pipeline (long rest periods between play bouts were not carried over, so most Not Playing time is unlabelled). Frames in which the calf is out of view are excluded, and Buck, Kick, Jump and Turn never decide a label. For other calves and frames, derive labels from the event logs and the frame index (example below). Time inside an annotated window with no coded state, and all time outside annotated windows, is unlabelled, not "Not Playing".

Tracking (tracking/<session_id>.tar): per calf and 3,000-frame folder, a JSON file {frame name: {"bounding_box": [x, y, w, h], "contour": [...]}}, and the initial prompt boxes (calf name: x, y, w, h). Tracked: Tol_2024-04-24 3 calves; Tol2_2024-06-11_AM/_PM 2 calves; Eem_2024-06-18 5 calves (folders 19–61); Eem_2024-06-19 calf 9. Tracks are automatic SAMURAI output and were not manually corrected; identity switches were not audited.

How to load

REPO is the repository ID (Sonam5/Calf-Play-Behavior-Dataset).

Download metadata and annotations only (a few MB):

python
from huggingface_hub import snapshot_download

REPO = "Sonam5/Calf-Play-Behavior-Dataset"
local = snapshot_download(REPO, repo_type="dataset",
                          allow_patterns=["README.md", "LICENSE", "metadata/*", "annotations/*"])

Download the frames of one session (e.g. 92 GB for Tol2_2024-06-11_PM):

python
snapshot_download(REPO, repo_type="dataset", allow_patterns=["frames/Tol2_2024-06-11_PM/*"], local_dir="calves")

Read one frame from a shard:

python
import io, tarfile
from PIL import Image
from huggingface_hub import hf_hub_download

path = hf_hub_download(REPO, "frames/Tol_2024-04-24/Tol_2024-04-24_video001.tar", repo_type="dataset")
with tarfile.open(path) as tf:
    img = Image.open(io.BytesIO(tf.extractfile("0000001.jpg").read()))

Stream shards with the webdataset library (each sample has the key 0000001 and the field jpg):

python
import webdataset as wds

urls = f"https://huggingface.co/datasets/{REPO}/resolve/main/frames/Tol_2024-04-24/Tol_2024-04-24_video{{001..030}}.tar"
for sample in wds.WebDataset(urls, shardshuffle=False).decode("pil"):
    key, img = sample["__key__"], sample["jpg"]

Join the labels with the frame index:

python
import pandas as pd

sid = "Tol_2024-04-24"
lab = pd.read_csv(f"{local}/annotations/labels/{sid}.csv.gz")
idx = pd.read_csv(f"{local}/metadata/frame_index/{sid}.csv.gz")
lab = lab.merge(idx[["frame_name", "chunk", "ocr_status"]], on="frame_name")

Label frames from the event logs (any calf, any aligned frame; example for one session):

python
sid = "Tol2_2024-06-12"
ev = pd.read_csv(f"{local}/annotations/event_logs/{sid}.csv")
st = ev[ev["Event_Type"] == "State start"].copy()
st["start"] = pd.to_datetime(st["observation_start"]) + pd.to_timedelta(st["Time_Relative_sf"], unit="s")
st["end"] = st["start"] + pd.to_timedelta(st["Duration_sf"], unit="s")

idx = pd.read_csv(f"{local}/metadata/frame_index/{sid}.csv.gz")
idx = idx[idx["ocr_status"] == "Success"].copy()
idx["t"] = pd.to_datetime(idx["ocr_timestamp"], errors="coerce") + pd.Timedelta(seconds=0.5)  # middle of the clock second
# screen OCR misreads before use (e.g. rolling-median check), then, per calf:
calf = st[st["calf_id"] == "TOL-2644"]
for _, e in calf.iterrows():   # states can overlap: Buck, Kick, Jump, Turn, Milk Feeding and Management are coded on top of the main state
    idx.loc[(idx["t"] >= e["start"]) & (idx["t"] < e["end"]), e["Behavior"]] = True

observation_start already contains the corrected start of every observation (for example the 24 observations of Eem_2024-06-20 whose typed start is missing or mistyped). Tol_2024-04-24 start times are 2 s earlier than the file names; see metadata/known_issues.md.

Usage notes

  • —Consecutive frames are nearly identical and several frames share one clock second. When training or evaluating models, split by blocks of time, by session or by farm rather than by random frames.
  • —The April and June 2024 Tol sessions share calves, and the two halves of 11 June 2024 are the same calves on the same day: treat all Tol sessions as one farm when comparing farms.
  • —The released frames, labels and tracks cover the annotated sessions and calves listed above. The videos of all 1,175 files are released, so users can decode further frames and run the OCR and tracking code (see Code availability) on other periods or calves.

Known issues (summary)

Read metadata/known_issues.md before use. The most important are:

  • —The April 2024 recording covers only about 05:40–24:00 daily; no video on 20–21 April 2024; no video 13:12:54–13:34:27 on 11 June 2024.
  • —Two videos cannot be played (no MP4 index): Tol_ch01_20240418-113301-121506.mp4 (corrupt) and Tol2_ch04_20240611_164646_172030.mp4 (truncated). They are released and flagged.
  • —Tol3 files spanning midnight are named by their end date and each contains a recording gap of about 30 min.
  • —Tol2_2024-06-11_AM frames 1–15,000 (06:09–06:59) were deleted; Eem_2024-06-18 has no frames for about 63 annotated minutes; the pilot day and Eem_2024-06-20 were not decoded. Two single frames are missing.
  • —OCR coverage is low for the 2024 Tol sessions and Success values include misreads (see above); 461 label rows (90 of the 452 rows of Tol2_2024-06-11_AM) sit on implausible readings.
  • —Event times use the position in the video file; because frames are dropped, they can differ from the camera clock by up to a few seconds towards the end of a file.
  • —Within the annotated windows, 2.8% of calf-time in Tol_2024-04-24 and 1.5% in Eem_2024-06-19 has no coded state (unlabelled).
  • —One observer per observation; inter-observer reliability was not measured.
  • —Channel labels differ between file names and overlays (Eem files ch04, overlay CH3; Tol2 annotation files say Ch01, camera is channel 4).
  • —Calves in neighbouring pens are visible in the Tol2 and Tol3 views; only the pen's own calves are annotated.

Not included

Recordings of the other farms of the observational study (no open-release agreement); leg-accelerometer data; per-calf image crops; decoded frames of two unannotated days; intermediate processing tables, which the event logs, labels and frame index supersede.

Ethics and consent

  • —The observational study was approved by the Animal Welfare Body of Utrecht University (approval 10818-2024-02). Recording was non-invasive.
  • —The farm owners consented to recording and to public release of these data under CC0 (no rights reserved).
  • —Farm staff occasionally appear in the footage; faces are not recognisable.
  • —Observer initials were replaced by O1–O5 in all annotation files. Farms are identified by the short codes Tol (with Tol2 and Tol3 denoting later recording rounds at the same farm) and Eem; no personal data of farmers are included.
  • —If you find identifiable personal information in the data, please contact us (below) so that it can be removed in a new revision.

Code availability

  • —https://github.com/Sonam525/Play-Behavior-Dataset: the code used to build this release (frame decoding, camera-clock OCR, tracking, labels and quality-control tables).
  • —https://github.com/Sonam525/Individual-Behavior-Analysis-with-CV: the video-analysis pipeline of the related play-behaviour study.

Licence

CC0 1.0 Universal. You may copy, modify and use the data for any purpose without asking permission. We ask, but do not require, that you cite the Data Descriptor and this dataset.

Citation

Until the Data Descriptor is published, please cite this dataset:

bibtex
@misc{calfplay_dataset_2026,
  author    = {Yang, Haiyu and Lesscher, Heidi and Liu, Enhong and Hostens, Miel},
  title     = {Video recordings and play behaviour annotations of group-housed dairy calves on two Dutch farms},
  year      = {2026},
  publisher = {Hugging Face},
  version   = {1.0},
  doi       = {10.57967/hf/10695},
  url       = {https://huggingface.co/datasets/Sonam5/Calf-Play-Behavior-Dataset}
}

Related preprints by the authors: Yang et al., arXiv:2602.00111 (2026); Yang et al., arXiv:2509.12047 (2025); Yang et al., arXiv:2603.17782 (2026).

Contact and maintenance

Corresponding author: Haiyu Yang, Department of Animal Science, Cornell University, Ithaca, NY, USA. Please open a discussion in the repository's Community tab.

Changelog

  • —Version 1.0 (2026-09-30): first complete release. The revision cited by the Data Descriptor is fixed under its DOI; later revisions get a new DOI and an entry here.