datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
git-commits
Dataset: dataset.jsonl
Auto-labeled commit dataset scraped from GitHub repositories. Each line is a JSON object representing one commit with extracted features and an inferred label.
Features
Field
Type
Description
Stats
text
string
Commit message first line, conventional prefix stripped
—
files_count
int
Number of files changed
mean 4.3, median 1, max 300
additions
int
Lines added
mean 88, median 6, max 187K
deletions
int
Lines deleted
mean 172… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/git-commits.synthetic-commit-msg-edits
✍️ Commit Message Edits Dataset - 🤖Synthetic
This dataset is a synthetic extension of our expert-labeled commit message edits dataset presented in Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings.
You can check Synthetic tab in our visualization app to browse through the datapoints!
Dataset Structure
Default
Default split contains the synthetic messages generated from expert-labeled dataset by an LLM.… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/synthetic-commit-msg-edits.openhands-commit-noise-databases
OpenHands Commit Noise Databases
This dataset contains commit-retrieval databases for 12 SWE-bench repositories at five noise ratios: 0%, 25%, 50%, 75%, and 100%.
Each archive expands to noise_NNN/<repository>/ directories containing:
commits.db: SQLite commit records
commits.faiss: normalized inner-product FAISS index
commits.meta.jsonl: FAISS row-to-commit metadata
commits.index_meta.json: embedding and index configuration
The 0% archive is an exact file-level copy of the… See the full description on the dataset page: https://huggingface.co/datasets/dengyixuan/openhands-commit-noise-databases.code-text-galeras-commit-generation-3k-dedupedbill_committees_us
Dataset Card for "bill_committees_us"
Dataset Summary
Dataset for US Congressional bills with committees information (bill_committees_us). Contains data for bills from the 108th to the 118th Congress, approximately 132,000 documents.
Supported Tasks and Leaderboards
More Information Needed
Languages
English
Dataset Structure
Data Instances
default
Data Fields
id: id of the bill in format(congress number +… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_committees_us.tw-ly-committee
Taiwan Legislative Yuan Committee Data(ly-tw-committee)
您也可以透過網頁介面瀏覽 committee 資料集:https://dataly.openfun.app/collection/list/committee
委員會基本資料比較少會變動,可視為參考資料(Reference Data)
Data Fields
資料欄位
說明
委員會代號
為 committee 的 id
委員會名稱
如欄位名稱所述
委員會職掌
為一段短文敘述該委員會的工作職責
委員會類別
int 總共三個類別: 1:常設委員會 2:特種委員會 3:國會改革前舊委員會名稱
委員會類別:str
類別的中文名稱
