THULab/phenology_normal_hawaii
Phenology-Normal-Hawaii (TsFile) This dataset is an Apache TsFile conversion of imageomics/phenology-normal-hawaii. Modalities: Time-series. Overview Vegetation color-index time series (GCC / RCC) from the PUUM (Pu'u Maka'ala) site, Hawaii, for fine-grained phenological analysis. Extracted from NEON PhenoCam images; daily curves with gcc_mean, gcc_50, rcc_*, midday_* etc. Each monitoring site is a device identified by the site TAG. Converted observations: 10… See the full description on the dataset page: https://huggingface.co/datasets/THULab/phenology_normal_hawaii.
Phenology-Normal-Hawaii (TsFile)
This dataset is an Apache TsFile conversion of `imageomics/phenology-normal-hawaii`.
Modalities: Time-series.
Overview
- Vegetation color-index time series (GCC / RCC) from the PUUM (Pu'u Maka'ala) site, Hawaii, for fine-grained phenological analysis.
- Extracted from NEON PhenoCam images; daily curves with
gcc_mean,gcc_50,rcc_*,midday_*etc. - Each monitoring site is a device identified by the
siteTAG.
- Converted observations: 10,945 rows across 1 TsFile file(s)
- Source format: csv
TsFile schema
- Time — source
date(datetime), converted to INT64 milliseconds.
Conversion notes
site(from the source file name) is a TAG so each site is a separate device.- Flag columns (
snow_flag,outlierflag_*) dropped as per-sample quality flags, not measurements.
Source & license
- Original dataset: https://huggingface.co/datasets/imageomics/phenology-normal-hawaii
- Author / publisher: imageomics
- License: cc-by-4.0
Usage
Install the Apache TsFile Python SDK (pip install tsfile) and read a converted file:
from pathlib import Path
from tsfile import TsFileReader
path = Path("phenology_normal_hawaii.tsfile")
with TsFileReader(str(path)) as reader:
schemas = reader.get_all_table_schemas()
print("tables:", list(schemas))
table_name = next(iter(schemas))
table = schemas[table_name]
columns = [column.get_column_name() for column in table.get_columns()]
print("columns:", columns)
field_names = [
column.get_column_name()
for column in table.get_columns()
if column.get_column_name() not in {"Time", "time"}
]
if field_names:
with reader.query_table(table_name, field_names[:3], batch_size=1024) as result:
batch = result.read_arrow_batch()
if batch is not None:
print(batch.to_pandas().head())