Collections
GiftEvalPretrain
GIFT-Eval Pre-training Datasets
Pretraining dataset aligned with GIFT-Eval that has 71 univariate and 17 multivariate datasets, spanning seven domains and 13 frequencies, totaling 4.5 million time series and 230 billion data points. Notably this collection of data has no leakage issue with the train/test split and can be used to pretrain foundation models that can be fairly evaluated on GIFT-Eval.
📄 Paper
🖥️ Code
📔 Blog Post
🏎️ Leader Board
Ethical Considerations… See the full description on the dataset page: https://huggingface.co/datasets/CollectionStudio/GiftEvalPretrain.Long-Data-Collections-Pretrain-Without-Books
Dataset Card for "Long-Data-Collections-Pretrain-Without-Books"
Paraquet version of the pretrain split of togethercomputer/Long-Data-Collections WITHOUT books
Statistics (in # of characters): total_len: 236088622215, average_len: 25159.041601590307
si_us_revolutionary_era_collections
Dataset Card for Smithsonian American Revolutionary Era Collections
Dataset Summary
A specially selected subset of the Smithsonian’s Open Access collections covering objects from 1770–1810 selected for the Revolution Crossroads project in honor of the 250th anniversary of the founding of the United States. Drawn from four museums—the National Museum of American History, National Postal Museum, Smithsonian American Art Museum, and National Portrait Gallery—the… See the full description on the dataset page: https://huggingface.co/datasets/RevolutionCrossroads/si_us_revolutionary_era_collections.rgc-collections
RGC analytical datasets
This dataset hosts derived Silver and Gold snapshots, analysis, model artifacts and analytical contracts. On 2026-10-03, the user requested keeping raw collection evidence in local files and clarified that only raw files should be removed from the current public tree.
The raw archive evidence/chocolate/uk.tar.gz, root raw index products.jsonl and root export-manifest.json are stored locally. Silver source-listings remain published, including embedded… See the full description on the dataset page: https://huggingface.co/datasets/CoralLeiCN/rgc-collections.speech-to-speech-collections
Real debt-collection calls (speech-to-speech format)
173 train / 3 validation calls from a production collections voice
agent (Hindi/Hinglish), processed into the same schema as
MeghanaKap/speech-to-speech-data.
Each call is stereo: channel 0 is the agent, channel 1 the customer, so both
sides can be modelled without speaker separation.
Privacy
Real customers are heard in these recordings, so access is gated and reviewed.
Kept: the customer's first name (the model… See the full description on the dataset page: https://huggingface.co/datasets/MeghanaKap/speech-to-speech-collections.3D_Room_Collections
3D Room Collections
这是多个室内场景数据集的字段导出集合(field-oriented exports)。每个数据集保存在自己的一级目录中,目录内包含该数据集的说明、提取脚本和 JSON 导出文件。主输出保存房间、物体、位置、尺寸、旋转以及源数据中已有的属性,不把 3D mesh、纹理或渲染结果混进统一 JSON。个别目录另保留了复现所需的源标注或辅助关键帧,具体范围以子目录 README 为准。
不同数据集的字段含义并不完全相同。例如,有些数据集提供真实 room boundary,有些只有 scan-level 或物体范围的代理边界;有些有 doors/windows,有些源数据没有这些标注。使用某个数据集前,请先阅读对应目录的 README.md。
目录结构 / Directory layout
对外发布时,保留根目录 README 和下面这些以 _exported 结尾的数据集目录即可:
3D_Room_Collections/
├── README.md
├──… See the full description on the dataset page: https://huggingface.co/datasets/imChuling/3D_Room_Collections.
