Team Ai
Datasetpublic

open-index/open-github-issues

OpenGitHub Issues What is it? The full development metadata of 7 public GitHub repositories, fetched from the GitHub REST API and GraphQL API, converted to Parquet and hosted here for easy access. Right now the archive has 6.0M rows across 8 tables in 699.6 MB of Zstd-compressed Parquet. Every issue, pull request, comment, code review, timeline event, file change, and CI status check is stored as a separate table you can load individually or query together. This… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-github-issues.

sourceHugging Faceodc-byupdated 4mo agoView on Hugging Face
1likes546downloads
stream_issues.py14 linesDownload Raw Back to code
1# /// script2# requires-python = ">=3.11"3# dependencies = ["datasets"]4# ///5"""Stream issues from the dataset without downloading everything."""6 7from datasets import load_dataset8 9ds = load_dataset("open-index/open-github-issues", "issues", streaming=True)10for i, row in enumerate(ds["train"]):11    print(f"#{row['number']}: [{row['state']}] {row['title']} (by {row['author']})")12    if i >= 19:13        break14