commitpackft
commitpackftCommitPackFT is is a 2GB filtered version of CommitPack to contain only high-quality commit messages that resemble natural language instructions.commitpackftCommitPackFT is is a 2GB filtered version of CommitPack to contain only high-quality commit messages that resemble natural language instructions.commitpack-ft-instruct-ratedThis is commitpack-ft-instruct, derived from Octocode's CommitPackFT, augmented with a quality analysis of the instruction-response pair by a local model. This did a pretty decent job of identifying pairs that obviously don't have enough context to know what change is being requested, or where the commit message does not match with the changes made.
Data files (yaml, plain text, json, etc.) were heavily downsampled in preparing this dataset to skew it more towards actual code work. All entries… See the full description on the dataset page: https://huggingface.co/datasets/chargoddard/commitpack-ft-instruct-rated.commitpack-ft-instructOctocode's CommitPackFT in Alpaca instruction format, with several randomly selected natural language preludes to the commit messages to make them better resemble a user request.
When the instruction, old code, and new code combined are small enough to fit within 4096 Llama tokens the output is usually the full contents of the file after a commit. Otherwise, the output will be a sequence of ndiff chunks with up to five lines of context each.
An example:
```ndiff
from… See the full description on the dataset page: https://huggingface.co/datasets/chargoddard/commitpack-ft-instruct.CommitPackFTcommitpackft
Dataset Card for CommitPackFT
Dataset Summary
CommitPackFT is a 2GB filtered version of CommitPack to contain only high-quality commit messages that resemble natural language instructions.
Creation: The dataset can be recreated using instructions available here.
Languages: 277
OctoPack🐙🎒:
Data
CommitPack
4TB of GitHub commits across 350 programming languages
CommitPackFT
Filtered version of CommitPack for high-quality commit messages that resemble… See the full description on the dataset page: https://huggingface.co/datasets/vnixxa31/commitpackft.
