Team Ai
Datasetpublic

THULab/github_top_developers

github_top_developers (TsFile) Apache TsFile version of diamond-in/github-top-developers. Records: 39,390 Schema (TsFile structure) name (TAG) — device dimension(s). name (FIELD). rank (FIELD). Usage Install the Apache TsFile Python SDK (pip install tsfile) and read a converted file: from pathlib import Path from tsfile import TsFileReader path = Path("github_top_developers.tsfile") with TsFileReader(str(path)) as reader: schemas =… See the full description on the dataset page: https://huggingface.co/datasets/THULab/github_top_developers.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes34downloads
README.md60 linesDownload Raw Back to root
1---2license: mit3task_categories:4- time-series-forecasting5tags:6- tsfile7- timeseries8- time-series9- format:tsfile10pretty_name: github_top_developers11size_categories:12- 10K<n<100K13---14 15# github_top_developers (TsFile)16 17Apache TsFile version of [`diamond-in/github-top-developers`](https://huggingface.co/datasets/diamond-in/github-top-developers).18 19- **Records:** 39,39020 21## Schema (TsFile structure)22 23- **name** (TAG) — device dimension(s).24- **name** (FIELD).25- **rank** (FIELD).26 27## Usage28 29Install the Apache TsFile Python SDK (`pip install tsfile`) and read a converted file:30 31```python32from pathlib import Path33from tsfile import TsFileReader34 35path = Path("github_top_developers.tsfile")36with TsFileReader(str(path)) as reader:37    schemas = reader.get_all_table_schemas()38    print("tables:", list(schemas))39    table_name = next(iter(schemas))40    table = schemas[table_name]41    columns = [column.get_column_name() for column in table.get_columns()]42    print("columns:", columns)43    field_names = [44        column.get_column_name()45        for column in table.get_columns()46        if column.get_column_name() not in {"Time", "time"}47    ]48    if field_names:49        with reader.query_table(table_name, field_names[:3], batch_size=1024) as result:50            batch = result.read_arrow_batch()51            if batch is not None:52                print(batch.to_pandas().head())53```54 55## Source & license56 57- Original dataset: https://huggingface.co/datasets/diamond-in/github-top-developers58- Author / publisher: diamond-in59- License: mit60