Team Ai
Datasetpublic

t-tech/T-ECD

T-ECD: T-Tech E-commerce Cross-Domain Dataset ⭐️ T-ECD is a large-scale anonymized cross-domain dataset for recommender systems research, created by T-Bank's RecSys R&D team. It captures real-world e-commerce interaction patterns across multiple domains while preserving privacy through a multi-stage anonymization pipeline. πŸ“„ Paper: T-ECD: A Large-Scale Cross-Domain E-Commerce Dataset for Industrial Recommender Systems β€” KDD '26 (32nd ACM SIGKDD Conference on Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/t-tech/T-ECD.

sourceHugging Facecc-by-nc-sa-4.0updated 5d agoView on Hugging Face
37likes11kdownloads
Dataset Card

T-ECD: T-Tech E-commerce Cross-Domain Dataset

image (2)

⭐️ T-ECD is a large-scale anonymized cross-domain dataset for recommender systems research, created by T-Bank's RecSys R&D team. It captures real-world e-commerce interaction patterns across multiple domains while preserving privacy through a multi-stage anonymization pipeline.

πŸ“„ Paper: T-ECD: A Large-Scale Cross-Domain E-Commerce Dataset for Industrial Recommender Systems β€” KDD '26 (32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining), Jeju Island, Republic of Korea. πŸ’» Baselines and reproducible pipeline: github.com/kliakhnovich/tecd-baselines

🎯 Overview T-ECD represents user interactions across five different e-commerce domains within a banking ecosystem:

  • β€”Marketplace β€” browsing and interacting with items in an e-commerce marketplace.
  • β€”Retail β€” interactions within a retail delivery service, including cart additions and completed orders.
  • β€”Payments β€” online and offline financial transactions between users and brands.
  • β€”Offers β€” responses to promotional content such as impressions, clicks, and partner transitions.
  • β€”Reviews β€” explicit user feedback in the form of ratings and embeddings of textual comments.

Scale:

  • β€”~135B interactions
  • β€”~44M users
  • β€”~30M items
  • β€”1300+ days of temporal coverage

Additionally, we provide T-ECD Small - a compact version containing 1B interactions that excludes the Payments domain.

<div style="font-size: 1.1em;">

MetricT-ECD SmallT-ECD Full
πŸ”„ Interactions~1B~135B
πŸ‘₯ Users~3.5M~44M
πŸ“¦ Items~2.6M~30M
πŸͺ Brands~29K~1M
πŸ“… Temporal Coverage200+ days1300+ days
🌐 Domains4 (excl. Payments)5 (all domains)

</div>

<img src="https://cdn-uploads.huggingface.co/production/uploads/645d4947f5760d1530d55023/Y3hHv_cipdq2p4A9jiQoz.png" style="max-width: 80%; height: auto;">

Cross-domain consistency is achieved by aligning identifiers across all domains:

  • β€”the same user_id always refers to the same individual user, and
  • β€”the same brand_id always refers to the same brand entity.

This alignment allows researchers to seamlessly link interactions from different services, enabling studies in transfer learning, cross-domain personalization, and multi-task modeling.

<img src="https://cdn-uploads.huggingface.co/production/uploads/645d4947f5760d1530d55023/QG0DavvcvccN1GcNgRL6.png" style="max-width: 80%; height: auto;"> <img src="https://cdn-uploads.huggingface.co/production/uploads/645d4947f5760d1530d55023/s8a8iC4RmUjsDhzOVPvD.png" style="max-width: 80%; height: auto;">


πŸ“‚ Data Schema

The dataset is stored in Parquet format with daily partitions ({day}). The directory structure is as follows:

t-ecd/
β”œβ”€β”€ users.pq
β”œβ”€β”€ brands.pq
β”œβ”€β”€ marketplace/
β”‚ β”œβ”€β”€ events/{day}.pq
β”‚ └── items.pq
β”œβ”€β”€ retail/
β”‚ β”œβ”€β”€ events/{day}.pq
β”‚ └── items.pq
β”œβ”€β”€ payments/
β”‚ β”œβ”€β”€ events/{day}.pq
β”‚ └── receipts/{day}.pq
β”œβ”€β”€ offers/
β”‚ β”œβ”€β”€ events/{day}.pq
β”‚ └── items.pq
└── reviews/{day}.pq
Data availability

<img src="https://cdn-uploads.huggingface.co/production/uploads/645d4947f5760d1530d55023/c2Clc9bNxL9i7jgGBfBq2.png" style="max-width: 80%; height: auto;" alt="Temporal distribution of events over domains"> Temporal distribution of events over domains In line with real-world industrial environments, domain-specific data availability varies in historical depth. This reflects practical constraints including data retention policies and product lifecycle stages - newer e-commerce services naturally have shorter histories compared to established banking domains like payments and transactions.

βš™οΈ Events and Catalogs

  • β€”Events: Each domain provides logs of user interactions with the following possible columns:
  • β€”action_type β€” interaction type (e.g., view, click, add-to-cart, order, transaction).
  • β€”subdomain β€” surface where the interaction occurred (recommendations, catalog, search, checkout, campaign); available in Marketplace and Retail.
  • β€”item_id β€” present in Marketplace, Retail, and Offers; identifies a specific product or offer.
  • β€”brand_id β€” present in all domains; denotes the seller, store, or partner associated with an item, offer, or transaction.
  • β€”price β€” represents the monetary value of the interaction.
  • β€”count β€” represents the amount of items in single interaction.
  • β€”os β€” user operating system, available in Marketplace and Retail.

<img src="https://cdn-uploads.huggingface.co/production/uploads/645d4947f5760d1530d55023/Q7aeb_I-Yf-rcqyPDTOLa.png" style="max-width: 80%; height: auto;" >

  • β€”Item catalogs (`items.pq`): Available for Marketplace, Retail, and Offers. Each entry includes:
  • β€”item_id
  • β€”brand_id
  • β€”category information (if available)
  • β€”pretrained embedding (if available)
  • β€”User catalog (`users.pq`): Contains anonymized user attributes such as region and socio-demographic cluster.
  • β€”Brand catalog (`brands.pq`): Contains brand_id, brand-level metadata, and embeddings.
🧾 Special Structures
  • β€”Receipts (`payments/receipts/{day}.pq`): Some transactions include detailed receipts with purchased items, their quantities, and prices. Items are aligned with Marketplace and Retail catalogs, enabling fine-grained cross-domain linkage at the product level.
  • β€”Reviews (`reviews/{day}.pq`): Provide explicit ratings per brand. Raw text reviews are not included; instead, we release pretrained text embeddings to preserve privacy while enabling multimodal research. ---

πŸ› οΈ Data Collection

T-ECD was generated through a multi-step process:

  1. 1.Sampling of event chains: sequences of interactions were sampled from real logs of T-Bank ecosystem services.
  2. 2.Anonymization: user and brand identifiers were pseudonymized; sensitive attributes removed.
  3. 3.Controlled perturbation: sampled chains were transformed with calibrated noise β€” timestamp jitter, perturbation of numerical attributes, generalization of rare values, and a small amount of event deletion paired with synthetic event injection β€” preserving structural properties such as sparsity, heavy tails, cross-domain overlaps, and behavioral contexts.

This process ensures that the dataset is privacy-preserving while remaining representative of industrial recommender system data.

πŸ“Š How much does anonymization cost?

We measured it instead of asking you to assume it. Using an identical preprocessing, splitting, and training pipeline on the original internal logs and on this released version, Recall@200 drops by 20–35% depending on domain and model β€” but the relative ranking of models is preserved. Conclusions you draw on T-ECD should therefore transfer. The full comparison across seven baselines and three domains is in the anonymization fidelity study (RQ1) of the paper.

⚠️ Important Note on Temporal Data Usage

<img src="https://cdn-uploads.huggingface.co/production/uploads/645d4947f5760d1530d55023/zaPAcuD3CItTzP2PBkErs.png" style="max-width: 80%; height: auto;">

To prevent data leakage, events from the final 12 hours should not be used for prediction tasks. The dataset contains temporal noise that requires maintaining a minimum 12-hour gap between the timestamp of the most recent user event and the prediction timestamp. This constraint applies to both training and testing scenarios to avoid temporal data leakage.

Download

Basic Download
python
from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="t-tech/T-ECD",
    repo_type="dataset",
    allow_patterns="dataset/full/",
    local_dir="./t_ecd_data",
    token="<your_hf_token>" 
)
Selective Download

For advanced usage including selection of domains and date ranges we provide custom downloader tecd_downloader.py

Example usage:

python
from tecd_downloader import download_dataset

download_dataset(
    token="<your_hf_token>",
    dataset_path="dataset/small",
    local_dir="t_ecd_small_partial",
    domains=["retail", "marketplace"],
    day_begin=1300,
    day_end=1308,
    max_workers=10
)

πŸ“š Citation

If you use T-ECD in your research, please cite the paper:

bibtex
@inproceedings{10.1145/3770855.3817588,
  author = {Liakhnovich, Kiryl and Nikiforova, Anna and Matveev, Nikita and Skurikhin, Maxim and Veselov, Ilya and Ananyeva, Marina},
  title = {T-ECD: A Large-Scale Cross-Domain E-Commerce Dataset for Industrial Recommender Systems},
  year = {2026},
  isbn = {9798400722592},
  publisher = {Association for Computing Machinery},
  address = {New York, NY, USA},
  url = {https://doi.org/10.1145/3770855.3817588},
  doi = {10.1145/3770855.3817588},
  booktitle = {Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2},
  pages = {9382–9391},
  numpages = {10},
  keywords = {recommender systems, large-scale dataset, benchmark dataset, cross-domain recommendation, e-commerce},
  location = {Republic of Korea},
  series = {KDD '26}
}

πŸ” License

This dataset is released under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0) licence