Team Ai
Datasetpublic

open-index/open-github-issues

OpenGitHub Issues What is it? The full development metadata of 7 public GitHub repositories, fetched from the GitHub REST API and GraphQL API, converted to Parquet and hosted here for easy access. Right now the archive has 6.0M rows across 8 tables in 699.6 MB of Zstd-compressed Parquet. Every issue, pull request, comment, code review, timeline event, file change, and CI status check is stored as a separate table you can load individually or query together. This… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-github-issues.

sourceHugging Faceodc-byupdated 4mo agoView on Hugging Face
1likes541downloads
load_repo.py20 linesDownload Raw Back to code
1# /// script2# requires-python = ">=3.11"3# dependencies = ["datasets"]4# ///5"""Load pull requests for a specific repo."""6 7from datasets import load_dataset8 9ds = load_dataset(10    "open-index/open-github-issues",11    "pull_requests",12    data_files="data/pull_requests/facebook/react/0.parquet",13)14df = ds["train"].to_pandas()15print(f"Loaded {len(df)} pull requests")16print(f"Merged: {df['merged'].sum()} ({df['merged'].mean()*100:.1f}%)")17print(f"\nTop 10 by lines changed:")18df["total_lines"] = df["additions"] + df["deletions"]19print(df.nlargest(10, "total_lines")[["number", "additions", "deletions", "total_lines"]].to_string(index=False))20