wangyz1999/X-EGO-CS
X-Ego-CS Ten players. One match. Ten simultaneous first-person recordings, each paired with a 64 Hz stream of that player's exact keyboard, mouse and view-angle inputs — all on a common, measured clock. Paper · Paper code · Collection pipeline Cross-Ego Demo (Pistol Round) Your browser cannot play this video — download it instead. All ten players' points of view, from the same pistol round, on one clock. Note: this demo concatenates the ten streams… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/X-EGO-CS.
X-Ego-CS
Ten players. One match. Ten simultaneous first-person recordings, each paired with a 64 Hz stream of that player's exact keyboard, mouse and view-angle inputs — all on a common, measured clock.
Paper · Paper code · Collection pipeline
Cross-Ego Demo (Pistol Round)
<video controls poster="https://huggingface.co/datasets/wangyz1999/X-EGO-CS/resolve/main/assets/multi-ego-sync-demo-pistol-poster.jpg" width="100%"> <source src="https://huggingface.co/datasets/wangyz1999/X-EGO-CS/resolve/main/assets/multi-ego-sync-demo-pistol-h264.mp4" type="video/mp4"> <source src="https://huggingface.co/datasets/wangyz1999/X-EGO-CS/resolve/main/multi-ego-sync-demo-pistol.mp4" type="video/mp4"> Your browser cannot play this video — <a href="https://huggingface.co/datasets/wangyz1999/X-EGO-CS/resolve/main/assets/multi-ego-sync-demo-pistol-h264.mp4">download it instead</a>. </video>
All ten players' points of view, from the same pistol round, on one clock.
Note: this demo concatenates the ten streams into a grid for display. The dataset itself ships them as individual per-player POV recordings — the grid is not a dataset artifact.
Action Data + Alignment Demo
<video controls poster="https://huggingface.co/datasets/wangyz1999/X-EGO-CS/resolve/main/assets/action-overlay-demo-poster.jpg" width="100%"> <source src="https://huggingface.co/datasets/wangyz1999/X-EGO-CS/resolve/main/assets/action-overlay-demo.mp4" type="video/mp4"> Your browser cannot play this video — <a href="https://huggingface.co/datasets/wangyz1999/X-EGO-CS/resolve/main/assets/action-overlay-demo.mp4">download it instead</a>. </video>
One player's POV with their actual inputs drawn on top, frame by frame: held keys light up, the mouse indicator tracks usercmd_mouse_dx/dy, and Δyaw/Δpitch read out the per-tick view change — all from the state_action/*.parquet trajectory for that clip.
The tick / time readout and the Offset control are the point of stage 06: video and ticks do not start together, so each clip carries a measured offset (here -0.73s) that pins a frame to the tick that produced it. Set it wrong and the inputs visibly desynchronise from the footage.
### 🆕 Updated — v2, September 2026 This release changes three things. In short: more recordings, tick-level action data for every clip, and the collection code itself. 1 — More recordings. 45 → 101 matches, now spanning 8 maps instead of Mirage only. 2,168 rounds, 21,680 ego-clips, 279.4 hours of synchronized first-person video. 2 — Tick-level action data, newly extracted. Every clip now ships a 64 Hzstate_action/*.parquettrajectory: position, health, view angles, per-tick view deltas, raw mouse counts, and 25 binary key columns decoded from the engine's input bitfield. Paired with a measured video-to-tick offset per clip inalign/*.json— read from the in-game HUD timer rather than assumed — so a frame can be matched to the exact tick that produced it. Coverage is 100% actions and 100% alignment across all 21,680 clips. 3 — The collection pipeline is open source. Everything here was produced by `multi-ego-cs`: an eight-stage, resumable pipeline from FACEIT match discovery through demo download, CS2 replay capture, action extraction, alignment, packaging and publishing. It is released so this dataset can be reproduced or extended to your own scale — see its scaling guide for what collecting 1000 matches actually costs. Migration note: the v1trajectory/*.csvtree has been replaced bystate_action/*.parquet. See Changes from v1.
Introduced in:
X-Ego: Acquiring Team-Level Tactical Situational Awareness via Cross-Egocentric Contrastive Video Representation Learning Yunzhe Wang, Soham Hans, Volkan Ustun University of Southern California, Institute for Creative Technologies (2025) arXiv:2510.19150
What makes this different
Most gameplay datasets give you one viewpoint, or many viewpoints that are not actually synchronised, or video with no ground-truth actions. This one gives all ten simultaneous egocentric views of the same match, each with the player's real inputs, aligned to a measured offset rather than an assumed one.
Clips are one player's life, not one round. A dead player's spectator camera follows their killer. Recording whole rounds would make a third of any "egocentric" corpus silently somebody else's point of view. Each clip here starts when the round goes live and ends at that player's death (or the round end).
Video and ticks are aligned by measurement. Screen capture starts an unpredictable fraction of a second after a demo resumes, so pairing frame i with tick start + i·(64/fps) drifts by hundreds of milliseconds — enough to put a flick on the wrong frame. Every round carries a per-player offset measured by reading the in-game HUD timer with template matching (digit confidence 0.98–0.99; the ten players in a round agree to ~50 ms).
Coverage is stated, not assumed. Every clip row carries has_video, has_actions and has_align.
Contents
video/<id>/<steamid>/round_<n>.mp4 first-person recording
state_action/match=<id>/round=<n>/<steamid>.parquet 64 Hz state + actions
align/match=<id>/round=<n>/offsets.json video ↔ tick offsets
metadata/<id>.json rounds, kills, alive windows
demo/<id>.dem raw CS2 replay
manifest/clips.csv one row per clip + coverage
manifest/rounds.csv, matches.csv, summary.json
match_round_partitioned.csv legacy indexTrajectory columns (one row per tick, 64 Hz)
Maps
Splits
Assigned per match — rounds within a match share players, economy and callouts, so splitting them across train and test would leak. Assignment hashes the match id, so a match keeps its split as the dataset grows.
Changes from v1
If you used an earlier version of this dataset, here is what moved:
trajectory/ has been removed — state_action/ covers every match it did and 56 more. Video paths are unchanged, and match_round_partitioned.csv keeps its columns but now indexes all 2,168 rounds instead of a subset.
Rather than hard-coding any of these paths, read them from the manifest: clips.csv carries video_path, actions_path and align_path per clip.
How to download
pip install --upgrade huggingface_hub
# Everything (~420 GB — see below for lighter options)
hf download wangyz1999/X-EGO-CS --repo-type dataset \
--local-dir ./X-EGO-CS --max-workers 8The full corpus is large because of the video. To pull only what you need:
# Manifests + action data + alignment, no video (~5 GB)
hf download wangyz1999/X-EGO-CS --repo-type dataset --local-dir ./X-EGO-CS \
--include "manifest/*" "metadata/*" "align/*" "state_action/*"
# One match's video
hf download wangyz1999/X-EGO-CS --repo-type dataset --local-dir ./X-EGO-CS \
--include "video/<match_id>/*"Usage
import json, polars as pl
from huggingface_hub import snapshot_download
root = snapshot_download("wangyz1999/X-EGO-CS", repo_type="dataset")
clips = pl.read_csv(f"{root}/manifest/clips.csv")
full = clips.filter(
(pl.col("has_video") == 1)
& (pl.col("has_actions") == 1)
& (pl.col("has_align") == 1)
)
row = full.row(0, named=True)
# NOTE: csv readers infer `steamid` as an integer, but it keys the offsets JSON
# as a *string*. Cast it, or the lookup raises KeyError.
steamid = str(row["steamid"])
# Paths come from the manifest, so the layout is never guessed.
video = f"{root}/{row['video_path']}"
traj = pl.read_parquet(f"{root}/{row['actions_path']}")
off = json.load(open(f"{root}/{row['align_path']}"))["players"][steamid]Aligning video with ticks
video_time = game_sec + offset_sec # the only sign convention used here
frame_index = round(video_time * off["video_fps"])offset_sec is typically negative (−0.7 to −3 s): the round goes live shortly before the capture's first frame. Two consequences worth handling:
- Early ticks map to negative frame indices. For a −0.768 s offset at 30 fps, tick 0 is frame −23. Most video libraries treat a negative index as counting from the end, so you would silently read the last frame of the clip. Drop or clamp frames below zero:
traj = traj.filter((pl.col("game_sec") + off["offset_sec"]) >= 0)- Use the per-player `offset_sec` for single-player work, and
round_offset_sec(the mean over non-outlier players) when compositing several views into one synchronised grid.
How it was collected
Collected with multi-ego-cs, an open end-to-end pipeline: discover matches via the FACEIT API → download demos → extract round and alive-window structure → drive CS2 to replay each demo and capture each player's viewpoint → extract 64 Hz state/action trajectories → align video to ticks by HUD-timer template matching → package and publish.
The pipeline is published so this corpus can be extended or reproduced at a different scale. Note that recording is real-time: roughly 45–60 minutes of capture per match, which is the dominant cost of any collection effort.
Limitations
- 1 of 21,680 clip has no measured video↔tick offset (the HUD timer never resolved cleanly). Check
has_alignper clip andusableper round rather than assuming. - Map distribution is skewed toward Mirage (46/101 matches); an unconstrained FACEIT sample reflects what players actually queue for.
- `usercmd_mouse_dx/dy` are raw device counts, not degrees. Converting needs per-player sensitivity, which demos do not record.
- Steam IDs are pseudonymous, not anonymous. They are real accounts.
- Recordings are 1280×720 (a small number at 1024×720) at ~30 fps, with HUD visible and spectator x-ray disabled.
Ethics and terms
Built from publicly available FACEIT match demos via FACEIT's own API and CDN. Users must comply with FACEIT's terms of service and Valve's terms for Counter-Strike 2. The footage shows real players' gameplay under pseudonymous Steam IDs; please respect their privacy expectations and do not attempt deanonymisation.
Citation
@article{wang2025x,
title={X-Ego: Acquiring Team-Level Tactical Situational Awareness via Cross-Egocentric Contrastive Video Representation Learning},
author={Wang, Yunzhe and Hans, Soham and Ustun, Volkan},
journal={arXiv preprint arXiv:2510.19150},
year={2025}
}License
MIT for the dataset annotations, manifests and pipeline code. Underlying gameplay footage and demo files originate from FACEIT-hosted matches.
