Sergasgr/codealign-commitpackft
CodeAlign — curated CommitPackFT (8 languages) Instruction/code pairs from bigcode/commitpackft, filtered for syntax validity (tree-sitter), per-language lint errors, cyclomatic complexity, internal duplication and cross-sample near-duplicates (MinHash/LSH). Built as the SFT set of CodeAlign; the pipeline, thresholds and the full curation report live in that repository. config rows contents sft (default) 122,018 accepted samples — the SFT training set minus the rows… See the full description on the dataset page: https://huggingface.co/datasets/Sergasgr/codealign-commitpackft.
CodeAlign — curated CommitPackFT (8 languages)
Instruction/code pairs from `bigcode/commitpackft`, filtered for syntax validity (tree-sitter), per-language lint errors, cyclomatic complexity, internal duplication and cross-sample near-duplicates (MinHash/LSH). Built as the SFT set of CodeAlign; the pipeline, thresholds and the full curation report live in that repository.
from datasets import load_dataset
sft = load_dataset("Sergasgr/codealign-commitpackft", "sft", split="train")Format
messages is ChatML (user prompt, assistant target file). Two prompt types, derived from the commit itself: new_file (write-from-spec; commit created the file) and edit (existing file + commit message as instruction). 84.6% of sft rows are edit.
Languages (sft)
Licensing and provenance
Only samples whose upstream license is one of apache-2.0, bsd-2-clause, bsd-3-clause, cc0-1.0, isc, mit, unlicense are included; the licence of every sample is in its license column and applies to that sample's code.
As in CommitPackFT, each row keeps its source commit and repos so the copyright holder can be identified (100.0% of sft rows carry provenance). Code authors who want their code removed can open an issue on the GitHub repository.
Known issues in the quality columns
lint_errors holds the values computed by the v1.0 curation run, which had linter bugs (the data was not re-curated):
- C++, JavaScript, TypeScript:
lint_errorsis always 0, so these languages were filtered on syntax, complexity and duplication only. cpplint's total was parsed from the wrong output stream (fixed in the repository afterwards); ESLint ≥ 9 rejects the--no-eslintrccommand line (still open). - Python: every value includes ruff's summary line (+1, or +2 when ruff also printed a fix hint), so the actual number of violations is 1–2 lower (fixed in the repository afterwards).
Privacy
- Rows containing a high-confidence secret (private key blocks, AWS / GitHub / Slack / Google / Stripe credentials) were dropped: 48 from
sft, 56 fromannotated. - E-mail addresses (other than documentation and GitHub no-reply domains) were replaced with
<EMAIL>in 8,272sftrows. - Detection is pattern-based and will miss some personal data; do not use this dataset to identify individuals.
Decontamination
13-gram overlap of every HumanEval and MBPP problem (prompt + canonical solution) against the Python samples is reported in src/notebooks/01_curation_report.ipynb of the GitHub repository.
Citation
Upstream data: Muennighoff et al., OctoPack: Instruction Tuning Code Large Language Models (2023), arXiv:2308.07124.
