Team Ai
Datasetpublic

Sergasgr/codealign-commitpackft

CodeAlign — curated CommitPackFT (8 languages) Instruction/code pairs from bigcode/commitpackft, filtered for syntax validity (tree-sitter), per-language lint errors, cyclomatic complexity, internal duplication and cross-sample near-duplicates (MinHash/LSH). Built as the SFT set of CodeAlign; the pipeline, thresholds and the full curation report live in that repository. config rows contents sft (default) 122,018 accepted samples — the SFT training set minus the rows… See the full description on the dataset page: https://huggingface.co/datasets/Sergasgr/codealign-commitpackft.

sourceHugging Faceotherupdated 15d agoView on Hugging Face
0likes23downloads
Dataset Card

CodeAlign — curated CommitPackFT (8 languages)

Instruction/code pairs from `bigcode/commitpackft`, filtered for syntax validity (tree-sitter), per-language lint errors, cyclomatic complexity, internal duplication and cross-sample near-duplicates (MinHash/LSH). Built as the SFT set of CodeAlign; the pipeline, thresholds and the full curation report live in that repository.

configrowscontents
sft (default)122,018accepted samples — the SFT training set minus the rows dropped for secrets (see Privacy)
annotated145,058every processed sample with its curation verdict (status, error)
python
from datasets import load_dataset
sft = load_dataset("Sergasgr/codealign-commitpackft", "sft", split="train")

Format

messages is ChatML (user prompt, assistant target file). Two prompt types, derived from the commit itself: new_file (write-from-spec; commit created the file) and edit (existing file + commit message as instruction). 84.6% of sft rows are edit.

Languages (sft)

languagerows
javascript45,084
python39,725
java15,779
c_sharp8,062
typescript4,339
cpp3,668
go3,034
rust2,327

Licensing and provenance

Only samples whose upstream license is one of apache-2.0, bsd-2-clause, bsd-3-clause, cc0-1.0, isc, mit, unlicense are included; the licence of every sample is in its license column and applies to that sample's code.

license`sft` rows
mit76,499
apache-2.026,849
bsd-3-clause11,913
bsd-2-clause3,887
isc1,423
unlicense985
cc0-1.0462

As in CommitPackFT, each row keeps its source commit and repos so the copyright holder can be identified (100.0% of sft rows carry provenance). Code authors who want their code removed can open an issue on the GitHub repository.

Known issues in the quality columns

lint_errors holds the values computed by the v1.0 curation run, which had linter bugs (the data was not re-curated):

  • —C++, JavaScript, TypeScript: lint_errors is always 0, so these languages were filtered on syntax, complexity and duplication only. cpplint's total was parsed from the wrong output stream (fixed in the repository afterwards); ESLint ≥ 9 rejects the --no-eslintrc command line (still open).
  • —Python: every value includes ruff's summary line (+1, or +2 when ruff also printed a fix hint), so the actual number of violations is 1–2 lower (fixed in the repository afterwards).

Privacy

  • —Rows containing a high-confidence secret (private key blocks, AWS / GitHub / Slack / Google / Stripe credentials) were dropped: 48 from sft, 56 from annotated.
  • —E-mail addresses (other than documentation and GitHub no-reply domains) were replaced with <EMAIL> in 8,272 sft rows.
  • —Detection is pattern-based and will miss some personal data; do not use this dataset to identify individuals.

Decontamination

13-gram overlap of every HumanEval and MBPP problem (prompt + canonical solution) against the Python samples is reported in src/notebooks/01_curation_report.ipynb of the GitHub repository.

Citation

Upstream data: Muennighoff et al., OctoPack: Instruction Tuning Code Large Language Models (2023), arXiv:2308.07124.