Team Ai
Datasetpublic

THULab/devops_predictive_logs

DevOps Predictive Logs Dataset (TsFile) Apache TsFile version of Snaseem2026/devops-predictive-logs. Overview A synthetic dataset of realistic DevOps log sequences for training and benchmarking predictive failure models. Each log entry describes one infrastructure observation (service, pod, level, message) inside one of 10 failure scenarios together with its incident metadata: severity, whether the pod eventually fails, and the time-to-failure in minutes. Rows:… See the full description on the dataset page: https://huggingface.co/datasets/THULab/devops_predictive_logs.

sourceHugging Facemitupdated 29d agoView on Hugging Face
0likes39downloads
Dataset Card

DevOps Predictive Logs Dataset (TsFile)

Apache TsFile version of `Snaseem2026/devops-predictive-logs`.

Overview

A synthetic dataset of realistic DevOps log sequences for training and benchmarking predictive failure models. Each log entry describes one infrastructure observation (service, pod, level, message) inside one of 10 failure scenarios together with its incident metadata: severity, whether the pod eventually fails, and the time-to-failure in minutes.

  • —Rows: 107 logs total — official train split 86 rows, test split 21 rows.
  • —Scenarios: 10 unique failure scenarios across 12 services / 13 pods.
  • —The repo additionally ships devops_logs_dataset.parquet, which is exactly the concatenation of the train and test splits (verified row-for-row), so only the two official splits are converted here.

Schema (TsFile structure)

Both splits share one schema; each split is its own .tsfile.

  • —Time (INT64, milliseconds) — log timestamp (%Y-%m-%d %H:%M:%S, naive).
  • —pod (TAG, STRING) — the pod emitting the log stream.
  • —level, service, message, scenario, issue_type, severity, dataset_version, created_date (FIELD, STRING)
  • —will_fail (FIELD, BOOLEAN)
  • —time_to_failure_minutes (FIELD, DOUBLE)

Usage

Install the Apache TsFile Python SDK (pip install tsfile) and read a converted file:

python
from pathlib import Path
from tsfile import TsFileReader

path = Path("devops_predictive_logs_test.tsfile")
with TsFileReader(str(path)) as reader:
    schemas = reader.get_all_table_schemas()
    print("tables:", list(schemas))
    table_name = next(iter(schemas))
    table = schemas[table_name]
    columns = [column.get_column_name() for column in table.get_columns()]
    print("columns:", columns)
    field_names = [
        column.get_column_name()
        for column in table.get_columns()
        if column.get_column_name() not in {"Time", "time"}
    ]
    if field_names:
        with reader.query_table(table_name, field_names[:3], batch_size=1024) as result:
            batch = result.read_arrow_batch()
            if batch is not None:
                print(batch.to_pandas().head())

Source & license

  • —Original dataset: <https://huggingface.co/datasets/Snaseem2026/devops-predictive-logs>
  • —License: MIT