ST-TraceWeaver/LLM-Inference-Traces-NYC
LLM-Inference-Traces-NYC A synthesized LLM inference workload for New York City: 2,015,645 requests, in 1,720,774 conversations, from 1,326,738 users, across 259 active geographic zones, over 7 consecutive days (2024-05-12 to 2024-05-18). Each request carries a user identifier, a timestamp, and a geographic zone — the three spatiotemporal attributes that public LLM conversation datasets are stripped of by privacy regulation. Without them, cache-aware scheduling, geo-distributed… See the full description on the dataset page: https://huggingface.co/datasets/ST-TraceWeaver/LLM-Inference-Traces-NYC.
LLM-Inference-Traces-NYC
A synthesized LLM inference workload for New York City: 2,015,645 requests, in 1,720,774 conversations, from 1,326,738 users, across 259 active geographic zones, over 7 consecutive days (2024-05-12 to 2024-05-18).
Each request carries a user identifier, a timestamp, and a geographic zone — the three spatiotemporal attributes that public LLM conversation datasets are stripped of by privacy regulation. Without them, cache-aware scheduling, geo-distributed load balancing, intelligent routing and temporal capacity planning cannot be evaluated on realistic traces.
This dataset is synthesized, not collected. User identifiers, timestamps and zones are fabricated by a synthesis engine. Nothing here is an observation of a real person, a real request, or a real location. Read Limitations before using it.
This repository was prepared for double-blind peer review: it carries no author names, affiliations or acknowledgements. It is otherwise a normal public release — the dataset, the code and the derived material are licensed for reuse as set out in `LICENSE`, subject to the third-party carve-outs described there.
### → Open the quick-look notebook Schema, data quality, workload shape, token distributions, user behaviour, inter-turn intervals and the NYC zone map — already executed, with every figure rendered. A minute to read, and by far the fastest way to see what is in the data.
How to load it
With `datasets`:
from datasets import load_dataset
ds = load_dataset("ST-TraceWeaver/LLM-Inference-Traces-NYC", split="train")
print(ds[0].keys())Directly from the parquet — the file is tracked with git-lfs, so clone with the LFS client installed or you will get a small pointer file instead of the data:
git lfs install
git clone https://huggingface.co/datasets/ST-TraceWeaver/LLM-Inference-Traces-NYC
cd LLM-Inference-Traces-NYC
git lfs pullimport polars as pl
df = pl.read_parquet("data/user_ts_zone_hotspot_combined.parquet")Read only what you need. The file is 642.3 MB, and almost all of it is the Conversation transcript column. The nine scalar columns load in a few seconds and about 165 MB:
import pyarrow.parquet as pq
scalars = ["UserID", "ConversationID", "Model", "Turn", "Timestamp",
"Zone", "HotspotID", "ContextToken", "GeneratedToken"]
df = pq.read_table("data/user_ts_zone_hotspot_combined.parquet", columns=scalars)Dataset at a glance
Every number in this table was measured directly from the released parquet — see Reproducing the statistics.
Traffic is dominated by one model. vicuna-13b accounts for 54.2% of requests, followed by koala-13b (7.3%), alpaca-13b (6.0%), vicuna-33b (3.3%) and llama-13b (2.9%). The set mixes hosted APIs (gpt-4, claude-2, palm-2) with open-weight models, so it spans both serving paths.
Nearly all users appear exactly once. 96.6% of users (1,281,823) have a single conversation, and each of those contributes exactly one request. That group still produces 63.6% of all requests, 69.0% of input tokens and 66.5% of generated tokens. The 3.4% of returning users carry the multi-turn structure.
The two populations are distinguishable by ID. The 1,281,823 single-request conversations use synthetic split_<n> conversation IDs. The other 733,822 rows carry hexadecimal IDs derived from the source chat data, and every incomplete row is in that group — so a null Model always means a real-conversation row that failed to finish, never a synthetic one.
Sessions are short-lived when they exist. Across the 105,456 conversations containing at least one non-zero gap between consecutive turns, the median gap is 31.0 minutes (p90 145.9, p99 505.6). 105,479 conversations have more than one turn; the gap statistic is slightly smaller because it discards turn pairs that share a timestamp.
Zone load is skewed. The median active zone serves 6,547 requests and the busiest serves 43,532.
Schema
ConversationID is not a uniform format, which matters if you group on it:
Two released columns are not among the nine attributes the paper declares for the synthesized request tuple $D{tgt}$: `ConversationID`, which is bookkeeping needed to group requests into conversations, and `HotspotID`, the Wi-Fi hotspot associated with the request — an additional spatial artifact produced alongside the zone. In the other direction, one paper attribute has no column of its own: $ci$, the number of turns in a conversation, is recoverable by grouping on ConversationID.
Which columns are synthesized
Three columns are fabricated: `UserID`, `Timestamp` and `Zone`. They are the product of the synthesis engine and correspond to nothing in the real world. Two more (ContextToken, GeneratedToken) are derived by tokenizing the conversation text. The rest is carried from upstream sources.
That provenance travels with the release in machine-readable form, so it survives the labels being copied out of this card:
- [`data/synthetic_attributes.json`](data/synthetic_attributes.json) — per-column provenance, method, cardinality, and a re-identification risk rating.
- The same column lists are mirrored in this card's YAML front matter under
synthetic_attributes. - The
provenance_vocabularyin the marker file defines the four values used:synthetic,derived,original,mixed.
import json
marker = json.load(open("data/synthetic_attributes.json"))
print(marker["synthesized_attribute_set"]) # ['UserID', 'Timestamp', 'Zone']
print(marker["columns"]["Zone"]["provenance"]) # 'synthetic'The three synthesized attributes are derived, not recorded, which is the point of the dataset:
- `UserID` is a cluster label, not an identity. Conversations are embedded and grouped with HDBSCAN into topically coherent partitions; conversations that cluster alone become the ephemeral users. The module explicitly does not claim to recover real user identities — no method could, since the source data is anonymised.
- `Timestamp` is DTW-aligned from a real inference trace onto each user's interleaved turns. The timestamp values are literal entries from the Azure trace; their pairing with conversations is synthetic.
- `Zone` is assigned by mobility-data fusion: empirical trip-based resampling blended with a per-minute population-weighted zone prior.
Provenance
Conversation content comes from LMSYS-Chat-1M. Request timestamps do not come from the conversation data — they are taken from the Azure Public Dataset LLM inference trace (2024) and aligned onto the conversation turns with dynamic time warping (DTW), so the inter-turn intervals reproduce the real gap distribution of a production inference service rather than the timing of the original chat collection. Geography is real NYC open data.
Core sources
These are the inputs that produced user_ts_zone_hotspot_combined.parquet.
data/sources/datacenters.csv has no public source URL recorded — it was compiled locally and is included here as-is.
LMSYS-Chat-1M and the Azure trace are not redistributed as files. The tabular NYC sources are included, so the spatial synthesis is reproducible from this repository alone, but rebuilding the combined parquet end to end still requires fetching those two upstream traces. Note that the Conversation column does reproduce LMSYS-Chat-1M text — that is a redistribution question in its own right, addressed under License.
How it was produced
These traces are reconstructed by ST-TraceWeaver, a configurable synthesis engine that reconstructs user identifiers, timestamps and geographic locations from anonymised conversation traces through three pluggable modules:
- Semantic Clustering — topically coherent user partitions, i.e. the synthesized user identifiers (HDBSCAN over conversation embeddings).
- DTW Timestamp Synthesis — behaviourally plausible request timestamps, from a cognitively motivated composite cost.
- Mobility-Data Fusion — geographic zone assignment by empirical trip-based resampling with population-weighted zone priors.
Each module is governed by tunable hyperparameters, so the pipeline can be re-run at a different operating point or on another trace. The engine is not published in this repository.
Synthesis fidelity (the paper's own measurements, not independently reproduced here): zone-level request volume preserves spatial rank order (Spearman ρ = 0.97; Pearson r = 0.86), and user activity is skewed with tunable concentration (Gini = 0.336, calibrated via the cluster size cap).
Intended uses
This dataset exists to evaluate systems that exploit the three spatiotemporal attributes. It is a workload for serving research, not a corpus for training or fine-tuning models.
- Cache-aware scheduling and personalized cache pre-warming
- Temporal load prediction and diurnal autoscaling
- Geo-distributed routing and data-sovereignty-compliant placement
- MoE inference optimization and model-affinity routing
- Workload-aware resource provisioning
- Reproducing or stress-testing the ST-TraceWeaver results at a different operating point
The configurable hyperparameter framework supports stress-testing under alternative behavioural assumptions, and the spatial module can be re-pointed at another city given three inputs — zone boundary definitions, a population distribution, and origin-destination trip records — with no algorithmic changes.
Not intended for: training or fine-tuning language models on the Conversation column; drawing conclusions about real people, real New Yorkers, or real request traffic; treating any UserID, Timestamp or Zone as an observation.
Limitations
The risk is the pairing, not the values
This is the limitation that matters most, and it cannot be designed away.
A single zone's conversations can be read as a neighbourhood profile. Filter this file to one Zone and you get the actual text of what people in that area asked an assistant about; group it by Timestamp and you get when they were active; follow a UserID and you get one person's requests linked across the week. None of the three columns is identifying on its own — a taxi-zone ID is a coarse areal unit, a timestamp is a point in time, an integer is an opaque label. The risk comes from their pairing.
That pairing is not a side effect of the release; it is the release. It is what makes cache-aware scheduling, geo-distributed routing and diurnal capacity planning evaluable at all. Coarsening Zone or quantizing Timestamp degrades exactly the workload properties the dataset exists to support, and does not remove the risk, because the conversations remain grouped and time-ordered either way. Any consumer who needs the utility is accepting the risk.
What each attribute actually is
- The zones come from taxi and mobility data — empirical trip-based resampling over NYC TLC trip records blended with population and MTA ridership priors. A
Zoneis where the model placed a request, not where anyone was. - The timestamps come from an inference trace — the Azure LLM Inference 2024 trace. They are real production arrival times, but they were produced by a different service in a different place and then DTW-aligned onto these conversations. They say nothing about when anyone actually typed.
- The user labels are cluster assignments, not real identities — HDBSCAN partitions of conversation embeddings. A
UserIDgroups requests that are topically similar. It is not a person, and the synthesis does not attempt to recover one. Two different real people with similar interests can share a label; one real person with diverse interests can be split across several. - The synthesized columns carry a machine-readable flag.
UserID,TimestampandZoneare markedsyntheticin `data/synthetic_attributes.json`, mirrored in this card's YAML front matter. That marker is what makes the labels survive being separated from this prose — if you redistribute or subset the release, carry it with you.
Other limitations
- Calibrated to New York City. The spatial synthesis is tuned to NYC mobility. Extending it to polycentric cities, car-dependent suburbs, or rural areas requires analogous mobility data and has not been empirically validated.
- Assumption-dependent. User identity synthesis assumes topical locality — that a user asks about related things over time — which need not hold for users with highly diverse interests. Spatial stationarity is a simplifying approximation.
- Parameter sensitivity. The DTW composite cost has five tunable weights whose optimal settings are deployment-specific, and the cognitive timing parameters require validation against LLM-specific interaction studies.
- No instance-level ground truth. Privacy regulations make individual user assignments unverifiable. Fidelity is argued at the distributional level only.
- `Timestamp` is timezone-naive, and
ZoneandHotspotIDare synthesized rather than observed. See Caveats. - The `Conversation` column is third-party text and may contain self-disclosed personal details written by members of the public. It is not licensed by the authors — see License.
Caveats
- Token caps are applied in this release.
ContextTokenis capped at 8,000 andGeneratedTokenat 1,500 per request, matching the source analysis and the paper; the released maxima equal those caps exactly. In the untruncated pipeline output 4,488 requests exceeded 8,000 input tokens (reaching 739,859) and 1,367 exceeded 1,500 generated (reaching 738,450). That tail is deliberately not published; re-derive it from the untruncated output if you need it. - Token counts track conversation length for typical rows — Pearson r ≈ 0.97 on log characters, at roughly 4 characters per token. The untruncated output contained a tail where this did not hold (the largest value,
ContextToken = 739,859, sat on under 4,000 characters of user text); capping at 8,000 / 1,500 removes that tail from this release, which is why the caps are applied rather than the tail published. - `Timestamp` carries no timezone. It is treated as UTC in the time-bucketed figures. This is an assumption, not a recorded fact.
- `HotspotID == -1` is a sentinel, not a hotspot — and it is 314,916 rows (15.6%, measured as 15.62%). Dropping it removes a substantial fraction of the data, so decide deliberately before filtering.
- 8,858 rows (0.44%) are incomplete.
Model,ContextToken,GeneratedTokenandConversationare null together, suggesting whole requests that failed to complete rather than random field dropout. Model-level aggregates include them as a null group. - Most datacenters have no type. Only 49 of 118 carry one (43 colocation, 6 building); the other 69 are drawn in grey as Unspecified. Zone request counts are also heavily skewed — the median zone serves ~6.5k against a maximum of ~43.5k — so the map's linear colour ramp leaves most zones in its light third.
- The map is a static image. It is built in Altair but displayed as a PNG and saved as a PDF, so it renders in any frontend without extra dependencies. Neither form carries the tooltips defined on the chart, so per-zone and per-hotspot values are not queryable in the output.
- `RequestCount` in the notebook's user-behaviour section counts distinct
(ConversationID, Turn)pairs, matching the source notebook — not raw row counts. - Inter-turn intervals are computed in timestamp order, not turn order, matching the notebook: the frame is sorted by
(ConversationID, Timestamp)and consecutive differences are taken per conversation, discarding non-positive gaps. Sorting byTurninstead gives a slightly different distribution (294,779 intervals in timestamp order versus 293,711 non-zero gaps in turn order).
Reproducing the statistics
Every number in Dataset at a glance and the figures in this card were measured from data/user_ts_zone_hotspot_combined.parquet with pyarrow 25.0.1 and polars 1.44, using the environment pinned in pyproject.toml and uv.lock. The quick-look notebook reproduces the analysis end to end:
uv sync
uv run jupyter nbconvert --to notebook --execute notebooks/quicklook.ipynb --output-dir outputsRelation to the paper
Where this artifact and the submitted paper differ, the paper is authoritative on method; the numbers here were measured from the released file rather than copied from the paper.
- Schema. The paper declares nine attributes for the synthesized request tuple $D_{tgt}$. The released parquet has ten top-level columns — see Schema for the column-to-symbol mapping and for the two extra columns (
ConversationID,HotspotID). - Token maxima. The parquet is capped at 8,000 input and 1,500 generated tokens per request, matching the paper's aggregate statistics exactly (
ContextTokenmax 8,000, sum 897,343,445;GeneratedTokenmax 1,500, sum 343,325,061). The untruncated pipeline output went higher — up to 739,859 and 738,450 — and that tail is deliberately not published; see Caveats.
License
This repository is not offered under a single license, and the card therefore sets license: other rather than claiming one it cannot substantiate. The full terms are in `LICENSE`. In short:
Why the parquet is not CC-BY-4.0 as a whole
data/user_ts_zone_hotspot_combined.parquet is one file whose columns have different upstreams. The authors can license the synthesized and derived columns, because those are their own work. They cannot license the Conversation column: it reproduces text from LMSYS-Chat-1M, which is distributed behind a gate under the LMSYS-Chat-1M Dataset License Agreement. That agreement grants a "limited, non-exclusive, non-transferable, non-sublicensable license to use the LMSYS-Chat-1M Dataset ... for both research and commercial purposes", and separately states:
Prohibited Transfers: You should not distribute, copy, disclose, assign, sublicense, embed, host, or otherwise transfer the dataset to any third party.
Source: the "LMSYS-Chat-1M Dataset License Agreement" section of <https://huggingface.co/datasets/lmsys/lmsys-chat-1m> (retrieved 2026-10-03). The dataset is gated — access requires accepting that agreement.
Practical consequences for users of this dataset:
- The authors grant no license to
Conversation,Model,Turnor the upstream-derivedConversationIDvalues, and cannot grant one. If you need the conversation text, obtain LMSYS-Chat-1M from the source and accept its terms yourself. - The upstream agreement is non-transferable, so nothing you receive from this repository gives you rights in that text.
- It also carries a non-identification obligation: users "must not attempt to identify the identities of individuals or infer any sensitive personal data encompassed in this dataset". The
UserIDvalues here are synthetic cluster labels and are not an attempt to recover real identities, but the obligation travels with the text. - The
Timestampvalues come from the Azure Public Dataset, which is published under CC-BY-4.0; attribution for those values belongs to Microsoft. - The NYC and MTA open-data tables carry their own attribution requirements, which govern reuse of the material they contributed.
Open item for the authors, recorded here deliberately: under the upstream "Prohibited Transfers" clause, redistributing the Conversation column publicly and ungated is not authorised by the LMSYS-Chat-1M terms. LICENSE section 3 sets out the three ways to resolve this — upstream permission, gating this repository behind the same agreement, or releasing the synthesized columns without the upstream text. Users and reviewers should treat the licensing of the Conversation column as unsettled until then.
Citation
If you use this dataset, please cite the ST-TraceWeaver paper. The full reference will be added on publication; until then, cite the artifact directly:
@misc{sttraceweaver_nyc_traces,
title = {LLM-Inference-Traces-NYC: a synthesized NYC LLM inference workload},
author = {{The ST-TraceWeaver authors}},
year = {2026},
note = {Dataset},
howpublished = {\url{https://huggingface.co/datasets/ST-TraceWeaver/LLM-Inference-Traces-NYC}}
}Please also cite the upstream sources whose data this artifact is built on:
- LMSYS-Chat-1M — Zheng et al., LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset, ICLR 2024. <https://arxiv.org/abs/2309.11998>
- Azure Public Dataset, LLM inference trace 2024 — the trace behind DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency, HPCA 2025. <https://github.com/Azure/AzurePublicDataset>
- NYC TLC Trip Record Data, MTA Subway Hourly Ridership, NYC Wi-Fi Hotspot Locations and the other NYC Open Data tables listed under Provenance.
Repository contents
Everything above a few kilobytes is stored through git-lfs. A clone made without git-lfs will contain small pointer files instead of the data — run git lfs pull to fetch the real bytes.
The notebook
notebooks/quicklook.ipynb is a first pass over the dataset. It needs the combined parquet plus the three small geo files in data/; it does not need the raw per-model traces or any of the upstream trip/ridership data.
It locates the parquet in this order:
$TRACEWEAVER_PARQUET, if set<repo>/data/user_ts_zone_hotspot_combined.parquet~/datasets/processed/user_ts_zone_hotspot_combined.parquet
To point it at a parquet somewhere else:
TRACEWEAVER_PARQUET=/path/to/user_ts_zone_hotspot_combined.parquet \
TRACEWEAVER_DATA_DIR=/path/to/geo/files \
uv run jupyter lab notebooks/quicklook.ipynbThe analysis frame loads only the nine scalar columns. The Conversation transcript column is inspected with a one-row peek of two columns (ConversationID + Conversation) rather than being loaded for all 2M rows, which keeps the working set at roughly 165 MB instead of the full file — small enough to run on a laptop. The notebook raises a clear error if it cannot find the parquet, and the geography section degrades gracefully if the geo files are absent (a clone that has not run git lfs pull will have pointer files for the two geo parquets).
Figures written to outputs/
All figures are PDF at 300 dpi with an embedded serif font stack.
The map is the one figure not drawn with matplotlib: it is built as Altair layers and exported through vl-convert-python, which is why both are dependencies. Its colours were checked with a palette validator — one single-hue sequential ramp for magnitude, plus two validated categorical hues for datacenter type and a neutral grey for the unlabelled ones.
Python environment (uv)
The environment is managed with uv and fully pinned by uv.lock.
curl -LsSf https://astral.sh/uv/install.sh | sh # once per machine
uv sync # create .venv from the lock
uv sync --locked # verify lock/pyproject consistency in CI--frozen is different and often confused with it: it means "never consult or rewrite uv.lock", so it installs whatever the lock currently says without noticing that pyproject.toml has moved on. Use it for air-gapped or read-only checkouts, not for drift detection. Plain uv sync re-resolves and may rewrite the lock.
uv add <package> # updates pyproject.toml and uv.lock together
uv remove <package>
uv lock --upgrade # refresh the lockfileCommit pyproject.toml and uv.lock together so everyone resolves identically.
Reproducibility
The figures committed in notebooks/quicklook.ipynb were produced from this repository's uv environment against the released parquet in data/.
