FA678945/github-issues-flask-augmented
🗂️ Dataset Card: GitHub Issues Dataset for Flask (Augmented) Author: FA678945 Version: 1.0 (2025) License: CC BY 4.0 Source: pallets/flask GitHub Repository 📘 Dataset Summary This dataset contains a curated collection of public GitHub Issues from the Flask web framework repository (pallets/flask). It has been augmented with derived metadata to support educational and research applications including sentiment analysis, text classification, and semantic search. Created for BSBDAT501, ICTDAT401… See the full description on the dataset page: https://huggingface.co/datasets/FA678945/github-issues-flask-augmented.
043
1---2configs:3- config_name: default4 data_files:5 - split: train6 path: data/train-*7dataset_info:8 features:9 - name: id10 dtype: int6411 - name: number12 dtype: int6413 - name: title14 dtype: string15 - name: user.login16 dtype: string17 - name: state18 dtype: string19 - name: comments20 dtype: int6421 - name: created_at22 dtype: string23 - name: updated_at24 dtype: string25 - name: labels26 list:27 - name: color28 dtype: string29 - name: default30 dtype: bool31 - name: description32 dtype: 'null'33 - name: id34 dtype: int6435 - name: name36 dtype: string37 - name: node_id38 dtype: string39 - name: url40 dtype: string41 - name: html_url42 dtype: string43 - name: body44 dtype: string45 - name: title_length46 dtype: int6447 - name: body_length48 dtype: int6449 - name: num_labels50 dtype: int6451 - name: is_bug52 dtype: bool53 - name: is_doc54 dtype: bool55 - name: is_feature_request56 dtype: bool57 splits:58 - name: train59 num_bytes: 1835660 num_examples: 1261 download_size: 2149562 dataset_size: 1835663---64🗂️ Dataset Card: GitHub Issues Dataset for Flask (Augmented)65 66Author: FA67894567Version: 1.0 (2025)68License: CC BY 4.069Source: pallets/flask GitHub Repository70 71📘 Dataset Summary72 73This dataset contains a curated collection of public GitHub Issues from the Flask web framework repository (pallets/flask).74It has been augmented with derived metadata to support educational and research applications including sentiment analysis, text classification, and semantic search.75 76Created for BSBDAT501, ICTDAT401, and BSBCRT404 assessment tasks, this dataset demonstrates ethical, compliant, and reproducible dataset creation using a public API.77 78📚 Supported Tasks & Benchmarks79 80This dataset is suited for:81 82Sentiment Analysis83 84Text Classification (bug vs non-bug, topic categorisation)85 86Semantic Search using FAISS87 88Keyword Extraction / Text Cleaning89 90EDA & NLP Demonstrations91 92Curriculum-Based Training for Dataset Creation and Documentation93 94This dataset is not intended for real-world decision-making or automated actions affecting individuals.95 96📡 Data Source97GitHub API Endpoint98https://api.github.com/repos/pallets/flask/issues99 100Collection Method101 102Data was retrieved using Python’s requests library:103 104import requests105 106url = "https://api.github.com/repos/pallets/flask/issues"107response = requests.get(url)108data = response.json()109 110 111Only public issues were collected.112No authentication, private repos, or sensitive data were involved.113 114🛠️ Data Fields115Raw Fields116Field Type Description117id int GitHub issue ID118title string Issue title119body string Issue body text120state string “open” or “closed”121labels list List of label objects122user string GitHub username (public)123created_at string ISO timestamp124updated_at string ISO timestamp125Augmented Fields126Field Type Description127title_length int Character count of title128issue_length int Character count of body129num_labels int Number of labels applied130is_bug bool True if issue text contains “bug”131created_month int Month extracted from timestamp132text_clean string Lowercased, punctuation-removed text133sentiment float/label Optional sentiment score/class134keyword_present bool Checks for keywords like “error” or “exception”135📁 Dataset Structure136 137Each record represents one GitHub Issue.138The dataset is provided as a structured table (e.g., JSON, CSV, Parquet, HF Dataset format).139 140🧪 Dataset Creation Process1411. Data Extraction142 143Retrieved via GitHub REST API144 145Limited to public issues146 147Includes only publicly visible usernames148 1492. Cleaning150 151Removed empty descriptions152 153Normalised whitespace154 155Lowercased and stripped punctuation156 1573. Augmentation158 159Added fields for length, sentiment, and keyword features to support NLP tasks.160 1614. Validation162 163Checked for malformed timestamps164 165Ensured all fields had consistent types166 167Verified no personal/sensitive data included168 169🔐 Ethical & Legal Considerations170Compliance with Policies171 172This dataset complies with:173 174Koorliny Kaatijin AI Use Policy175 176Only public data used177 178Full transparency and documentation179 180Data Access & Classification Policy181 182Classified as Public183 184Privacy Policy185 186No personal or sensitive information collected187 188Information Security & Acceptable Use Policy189 190Stored on an approved cloud repository (Hugging Face)191 192Data Retention & Disposal Policy193 194Dataset may be removed after academic use195 196Privacy197 198This dataset uses only publicly available GitHub usernames, which are explicitly non-sensitive.199 200No email addresses, private repositories, or personal data were collected.201 202Copyright & Licensing203 204Original GitHub content is subject to GitHub Terms of Service.205 206Augmented dataset is released under CC BY 4.0.207 208Attribution is required (see below).209 210⚠️ Limitations211 212Informal writing style in issues213 214Small sample size215 216Heuristic fields like is_bug may be noisy217 218Label and topic imbalance219 220Possible contributor-style bias221 222Not suitable for production ML systems.223 224🌐 Intended Uses225 226This dataset is intended for:227 228Classroom/academic demonstrations229 230Text preprocessing and NLP tutorials231 232Training on dataset creation & documentation processes233 234Experimenting with FAISS search235 236Exploratory data analysis237 238Not suited for:239 240Automated decision-making241 242High-risk AI applications243 244Profiling individuals245 246📜 Licensing247Dataset License248 249Creative Commons Attribution 4.0 (CC BY 4.0)250 251Attribution Requirements252 253Please credit:254 255Original Source Data:256pallets/flask GitHub repository257 258Issue Authors & Contributors259 260Dataset Creator: FA678945 (2025)261 262Dataset URL: Hugging Face link263 264🧾 Citation265FA678945 (2025). GitHub Issues Dataset for Flask (Augmented).266Hugging Face Datasets. Source data from the pallets/flask GitHub repository.267Licensed under CC BY 4.0.268 269📈 Future Work270 271Possible enhancements:272 273Embeddings for FAISS (Task 3)274 275Larger set of issues276 277Label-based classification tasks278 279Triage prediction model280 281Topic modelling