tech
Datasets
All datasets matching “tech”techbasetech-articles-live持续收集国内各个科技公众号的文章并提供简易分析工具,目前工具位于新智元公众号文件夹下。
使用的自己修改的微信公众号文章导出器,项目链接:我的GitHub仓库链接
T-ECD
T-ECD: T-Tech E-commerce Cross-Domain Dataset
⭐️ T-ECD is a large-scale anonymized cross-domain dataset for recommender systems research, created by T-Bank's RecSys R&D team.
It captures real-world e-commerce interaction patterns across multiple domains while preserving privacy through a multi-stage anonymization pipeline.
📄 Paper: T-ECD: A Large-Scale Cross-Domain E-Commerce Dataset for Industrial Recommender Systems — KDD '26 (32nd ACM SIGKDD Conference on Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/t-tech/T-ECD.wm_imagined
Imagined Data
This repository hosts imagined interaction data generated by world models across different environments, tasks, and data sources. Data are organized into separate subdatasets, with additional types of imagined data to be added over time.
The repository contains the RoboTwin2.0 and RealWorld subdatasets. Storage formats, field definitions, and state and action semantics are documented separately for each subdataset.
Dataset Index
Subdataset… See the full description on the dataset page: https://huggingface.co/datasets/Cirquar-Tech/wm_imagined.IndustryCorpus2_technology_scientific_research
IndustryCorpus2: Technology & Research
This repository contains the IndustryCorpus2: Technology & Research domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_technology_scientific_research.quranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.
