Sulak2020/github-issues-multirepo-example
Multi-repo GitHub Issues — example dataset Short descriptionThis dataset contains GitHub issues collected from multiple large open-source repositories (sampled): huggingface/datasets, pytorch/pytorch, tensorflow/tensorflow. Each row is one GitHub Issue (pull requests excluded). The dataset includes derived fields (time_to_close_days, label_count, body lengths) suitable for analytics and demonstration of dataset publishing workflows. Files data/github_issues_multirepo.csv — CSV… See the full description on the dataset page: https://huggingface.co/datasets/Sulak2020/github-issues-multirepo-example.
033
1---2language: en3license: cc-by-4.04tags:5 - analytics6 - software7 - github-issues8 - small-to-medium9---10 11# Multi-repo GitHub Issues — example dataset12 13**Short description** 14This dataset contains GitHub issues collected from multiple large open-source repositories (sampled): `huggingface/datasets, pytorch/pytorch, tensorflow/tensorflow`. Each row is one GitHub Issue (pull requests excluded). The dataset includes derived fields (time_to_close_days, label_count, body lengths) suitable for analytics and demonstration of dataset publishing workflows.15 16**Files**17- `data/github_issues_multirepo.csv` — CSV table.18- `data/github_issues_multirepo.parquet` — Parquet file.19- `data/github_issues_multirepo.jsonl` — JSON Lines.20 21**Schema (columns)**: repo, issue_id, number, title, user_login, state, created_at, closed_at, time_to_close_days, comments, label_count, labels, has_labels, body_char_count, body_word_count, html_url, comments_url, curator_comment22 23**How data was collected**24Data were retrieved via the GitHub REST API v3 (https://api.github.com) by paging the issues endpoints for each repository listed above. Collection date: 2025-10-25. Some fields (e.g., `sample_comments`) may have been omitted to reduce sensitive content.25 26**License, privacy & attribution**27- This derived dataset is released under **CC-BY-4.0**.28- Source data are public GitHub issues; original issue authors retain their original contributions under GitHub's terms. Each record contains `html_url` linking to the original issue — please attribute as appropriate.29- The dataset includes public usernames. If you will publish models or downstream artifacts, consider pseudonymizing or removing `user_login` and `sample_comments`.30- Do not assume personal data minimization: if your use case requires removal of personal data or compliance with local laws (e.g., GDPR), perform appropriate redaction and document it here.31 32**Intended uses and limitations**33- For exploratory analysis, project analytics, and demonstration of dataset publishing. Not intended as a comprehensive snapshot of repository history; rate limits and pagination caps may have truncated older issues.34 35**Contact / removal requests**36If an author requests removal of a specific record, please provide the issue URL and we will remove it from this dataset build. (Provide contact info or link to hub profile.)37 38 