OpenTSLab/SciTS
SciTS: Scientific Time Series Understanding and Generation with LLMs This repository contains the official dataset for SciTS: Scientific Time Series Understanding and Generation with LLMs (ICLR 2026). SciTS is a large-scale benchmark designed to evaluate the capabilities of large language models on complex scientific time series data. It spans 12 scientific disciplines, 43 distinct tasks, and includes 54,023 instances. Dataset Structure The benchmark is organized… See the full description on the dataset page: https://huggingface.co/datasets/OpenTSLab/SciTS.
53.5k
1---2configs:3- config_name: default4 data_files:5 - split: test6 path: meta_data.jsonl7license: cc-by-nc-sa-4.08task_categories:9- time-series-forecasting10- question-answering11language:12- en13tags:14- time series15- timeseries16- audio17- benchmark18- time series Reasoning19- time series Classification20- time series QA21- time series Anomaly Detection22- classification23- anomaly detection24size_categories:25- 10K<n<100K26pretty_name: 'SciTS: Scientific Time Series Understanding and Generation with LLMs'27---28 29# SciTS: Scientific Time Series Understanding and Generation with LLMs30 31 32This repository contains the official dataset for [**SciTS: Scientific Time Series Understanding and Generation with LLMs** (ICLR 2026)](https://openreview.net/forum?id=5YXccEP6uc). SciTS is a large-scale benchmark designed to evaluate the capabilities of large language models on complex scientific time series data. It spans 12 scientific disciplines, 43 distinct tasks, and includes 54,023 instances.33 3435 36## Dataset Structure37 38The benchmark is organized into a main `meta_data.jsonl` file, a `process` directory for handling restricted datasets, and 38 individual dataset folders. Each folder is named using the convention: `Domain-DatasetName-Scene-Task`.39 40```41├── process/42│ ├── process_ETT.py43│ ├── process_iNaturalist.py44│ ├── infer_template.py45│ ├── eval.py46│ └── requirements.txt47├── Domain-DatasetName-Scene-Task_1/48│ ├── raw_input_data/49│ └── raw_gt_data/ (for generation tasks)50├── Domain-DatasetName-Scene-Task_2/51│ └── raw_input_data/52...53├── Domain-DatasetName-Scene-Task_38/54│ ├── raw_input_data/55│ └── raw_gt_data/56└── meta_data.jsonl57```58 59- **`process/`**: Contains utility scripts, including `process_ETT.py` and `process_iNaturalist.py` for processing restricted datasets which cannot be released directly due to license restrictions, `infer_template.py` as an inference template, `eval.py` for evaluation, and `requirements.txt` for dependency installation.60- **`Dataset Folders`**: Each of the 38 folders contains the raw time series data for a specific dataset. `raw_input_data` holds the input signals, while `raw_gt_data` (present only for generation tasks) holds the ground truth output signals.61- **`meta_data.jsonl`**: A JSON Lines file containing metadata for every instance in the benchmark. Each line corresponds to one data sample.62 63### Dataset Collection64 65The 38 released datasets are listed below:66 67| Domain | Dataset Folder Name | Task ID |68| :--- | :--- | :--- |69| Astronomy | `Astronomy-GWOSC_GW_Event-Gravitational_wave-Anomaly_detection+Event_localisation` | ASU01, ASG02 |70| | `Astronomy-LEAVES-Light_curve-Classification` | ASU03 |71| Earth Science | `Earth_Science-STEAD-Earthquake-Anomaly_detection+Event_localisation` | EAU01, EAG02 |72| Bioacoustics | `Bioacoustics-Powdermill-Birds_vocalisation-Classification` | BIU01 |73| | `Bioacoustics-MarmAudio-Marmoset_vocalisation-Classification` | BIU03 |74| Meteorology | `Meteorology-TS_MQA-Weather-Anomaly_detection` | MEU01 |75| | `Meteorology-TIMECAP-Rainfall-Anomaly_detection` | MEU02 |76| | `Meteorology-MT_bench-Temperature-Forecasting` | MEG03 |77| | `Meteorology-MT_bench-Temperature-MCQ` | MEU04 |78| Economics | `Economics-FinMultiTime-Stock_closing_price-Forecasting` | ECG01 |79| | `Economics-MT_bench-Stock_price-Forecasting` | ECG02 |80| | `Economics-MT_bench-Stock-MCQ` | ECU03 |81| Neuroscience | `Neuroscience-MDD-Depressive_disorder-Anomaly_detection` | NEU01 |82| | `Neuroscience-TUEV-EEG_pattern-Classification` | NEU02 |83| | `Neuroscience-TS_MQA-EEG_signal-Forecasting` | NEG03 |84| | `Neuroscience-TS_MQA-EEG_signal-Imputation` | NEG04 |85| | `Neuroscience-WBCIC_SHU-Motor_imagery-Classification` | NEU05 |86| | `Neuroscience-Sleep-Sleep_staging-Classification` | NEU06 |87| Energy | `Energy-NewsForecast-Electronic_load-Forecasting` | ENG01 |88| | `Energy-TextETT-Sensor_signal_trend-Synthesis` | ENG03 |89| | `Energy-TS_MQA-Comprehensive_electricity-Forecasting` | ENG04 |90| | `Energy-TS_MQA-Comprehensive_electricity-Imputation` | ENG05 |91| Physiology | `Physiology-PTB_XL-ECG_status-Classification` | PHU01 |92| | `Physiology-TS_MQA-Physiological_signal-Forecasting` | PHG02 |93| | `Physiology-TS_MQA-Physiological_signal-Imputation` | PHG03 |94| | `Physiology-TS_MQA-ECG-Anomaly_detection` | PHU04 |95| | `Physiology-TS_MQA-Gait_freezing-Anomaly_detection` | PHU05 |96| | `Physiology-TS_MQA-Human_activity-Classification` | PHU06 |97| Urbanism | `Urbanism-NewsForecast-Traffic_flow-Forecasting` | URG01 |98| | `Urbanism-TS_MQA-Pedestrian_flow-Forecasting` | URG02 |99| | `Urbanism-TS_MQA-Pedestrian_flow-Imputation` | URG03 |100| | `Urbanism-TS_MQA-Traffic_flow-Anomaly_detection` | URU04 |101| | `Urbanism-MetroTraffic-Traffic_volume-Forecasting` | URG05 |102| Manufacturing | `Manufacturing-CWRU-Bearings_fault_location+Bearings_fault_size-Classification` | MFU01, MFU02 |103| | `Manufacturing-MIMII_Due-Machine_malfunction-Anomaly_detection` | MFU03 |104| Radar | `Radar-RadSeg-Coding_scheme-Classification` | RAU01 |105| | `Radar-RadarCom-Modes_and_modulation-Classification` | RAU02 |106| Math | `Math-Chaotic-Chaotic_system-Forecasting` | MAG01 |107 108## `meta_data.jsonl` Format109 110Each line in this file is a JSON object with the following structure, providing all necessary metadata to load and use a data sample.111 112```json113{114 "task_id": ["TASK_ID"], // List of task IDs associated with this sample (e.g., ["ASU03"] or ["ASU01", "ASG02"] for merged datasets)115 "id": "DATASET_ID", // Unique identifier of this sample within the dataset116 "data_type": "csv"/"npy"/"wav"/"flac", // File format of the raw time series data117 "input_ts":{118 "num_channel": int, // Number of channels (dimensions) in the input signal119 "channel_detail": [], // List of channel names, empty if none120 "path": "raw_input_data/sample_001_input.npy",121 "length": int, // Length of the input time series122 "timestamps": [], // Auxiliary timestamp information, empty if none123 "fs": int // Sampling frequency in Hz124 },125 "input_text": "INPUT_TEXT", // Textual prompt or task instruction provided as input126 "gt_text": "GT_TEXT", // Ground truth textual answer (for understanding tasks; empty for generation tasks)127 "gt_ts": {128 "path": "raw_gt_data/sample_001_output.npy",129 "length": int // Length of the ground truth time series130 },131 "gt_result": { ... }, // Structured ground truth result; format varies by task type (see below)132 "meta_data": {} // Additional metadata from the original data source133}134```135 136### `gt_result` Field Format137 138The structure of the `gt_result` field varies depending on the task type. This field provides the original ground truth for metric computation.139 140**1. MCQ**141```json142"gt_result": {143 "answer": "TEXT" // The correct textual answer144}145```146 147**2. Synthesis, Forecasting, Imputation**148```json149"gt_result": {150 "num_channel": int, // Number of channels (dimensions) in the ground truth signal151 "channel_detail": [], // List of channel names, empty if none152 "timestamps": [] // Auxiliary timestamp information, empty if none153}154```155 156**3. Classification**157 158For the `CWRU` dataset, which involves two classification sub-tasks, the category keys in class_list and gt_class are `"diameter"` and `"position"` respectively. For all other classification tasks, the category key is `"default"`.159 160```json161"gt_result": {162 "class_list": {163 "default": ["class_A", "class_B"], // List of candidate classes for each category164 ...165 },166 "gt_class": {167 "default": ["GT_CLASS"], // Ground truth class label for each category168 ...169 }170}171```172 173**4. Anomaly Detection**174```json175"gt_result": {176 "contain": Boolean // Boolean indicating if the required event is present177}178```179 180**5. Anomaly Detection + Event Localisation**181 182For the `GWOSC GW Event` and `STEAD` datasets, each of which includes both an `Anomaly Detection` task and an `Event Localisation` task, the gt_result field is defined in the following combined format:183 184```json185"gt_result": {186 "contain": Boolean, // Boolean indicating if the required event is present187 "start_time": int // The event index if contain is true, else null188}189```190 191## Handling Restricted Datasets192 193Due to license restrictions, the **ETT** (`ENG02`) and **iNaturalist** (`BIU02`) datasets are not directly included in this repository. To use them, the user need to download the original data and run the provided processing scripts.194 195**Step 1: Download the Data**196 197- **ETT**: Download `ETTh1.csv` from the official repository: [https://github.com/zhouhaoyi/ETDataset](https://github.com/zhouhaoyi/ETDataset)198- **iNaturalist**: Download the `Test Recordings` from the official repository: [https://github.com/visipedia/inat_sounds/tree/main/2024](https://github.com/visipedia/inat_sounds/tree/main/2024)199 200**Step 2: Install Dependencies**201 202Before running the processing scripts, install the required Python packages:203 204```shell205pip install -r process/requirements.txt206```207 208**Step 3: Run the Processing Script**209 210Place the downloaded files into a local directory. Then, from the root of this repository, run the corresponding script to process the data into the standard benchmark format.211 212- For ETT:213 ```shell214 python process/process_ETT.py --data_path /path/to/your/ETTh1.csv215 ```216 217- For iNaturalist:218 ```shell219 python process/process_iNaturalist.py --data_folder /path/to/your/iNaturalist/test220 ```221 222This will generate the `Energy-ETT-Transformer_sensor_signal-Forecasting` and `Bioacoustics-INaturalist-Animal_vocalisation-Classification` folders along with their `raw_input_data`, `raw_gt_data` subdirectories, as well as the processed test files.223 224## Baseline Inference and Evaluation225 226The `process` directory also includes scripts for running inference and evaluating the results.227 228### Inference229 230`process/infer_template.py`: Template code for the inference script. Implement the `initialize_model` function, then inference can be done by running:231 232```shell233python process/infer_template.py --scits_dir /path/to/scits_dir --output_dir /path/to/output_dir234```235 236### Evaluation237 238`process/eval.py`: Evaluation script. Run:239 240```shell241python process/eval.py evaluate --infer_dir /path/to/infer_dir242```243 244The evaluation results will be saved to `/path/to/infer_dir/results/`.245 246## Citation247 248If you use the SciTS benchmark, please cite the paper:249 250```bibtex251@inproceedings{252 wu2026scits,253 title={Sci{TS}: {S}cientific Time Series Understanding and Generation with {LLM}s},254 author={Wen Wu and Ziyang Zhang and Liwei Liu and Xuenan Xu and Jimin Zhuang and Ke Fan and Qitan Lv and Junlin Liu and Chen Zhang and Zheqi Yuan and Siyuan Hou and Tianyi Lin and Kai Chen and Bowen Zhou and Chao Zhang},255 booktitle={The Fourteenth International Conference on Learning Representations},256 year={2026},257 url={https://openreview.net/forum?id=5YXccEP6uc}258}259```