Team Ai
Datasetpublic

AISE-TUDelft/multilingual-code-comments-fixed-8

Fixed-8 Based on fixed-7 revision 14e85fe00a8b284cd226c58281ddd8e6b990b190. Replaces six Greek rows with missing expert labels with six newly labelled samples. All five language configurations retain 500 training rows (2,500 total). All other rows are unchanged. Removed ID Replacement ID 8000_5 1056_0 8000_14 4357_8 8000_15 4848_9 8000_16 29069_13 8000_17 1385_4 8000_18 5142_0 All 500 Greek rows now have all five expert accuracy labels. Original… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/multilingual-code-comments-fixed-8.

sourceHugging Faceupdated 6d agoView on Hugging Face
0likes138downloads
Dataset Card

Fixed-8

Based on fixed-7 revision 14e85fe00a8b284cd226c58281ddd8e6b990b190. Replaces six Greek rows with missing expert labels with six newly labelled samples. All five language configurations retain 500 training rows (2,500 total). All other rows are unchanged.

Removed IDReplacement ID
8000_51056_0
8000_144357_8
8000_154848_9
8000_1629069_13
8000_171385_4
8000_185142_0

All 500 Greek rows now have all five expert accuracy labels. Original supplied predictions and error-code strings are preserved; blank error-code fields are retained as provided. This release does not remove pre-existing file reuse or certify absence of partial code clones. Scores and judge results from fixed-7 must be recomputed for the six replacement IDs.

CodeGemma comment extraction correction

Corrected from dataset revision 292880b74aa2ec58dc1b00524a70b060359bf0fa. Only predicted_comment_google/codegemma-7b was changed: each affected completion is truncated immediately before its first <|file_separator|> token. The token and all subsequent generated text are removed; the preceding text and whitespace are preserved exactly. Raw predict_google/codegemma-7b generations, prompts, reference comments, expert labels, error codes, row order, and every other generator's columns are unchanged.

LanguageCorrected comments
Chinese464
Dutch480
English477
Greek469
Polish469

11 affected completions contain only whitespace or no text before the boundary; these are preserved as such, with their supplied expert labels unchanged. Previously computed neural scores for changed comments must be recomputed from these corrected inputs.

Remaining comment extraction corrections

Corrected from revision d02cf177485908c0d26e71c26f4a584c0e38dd5d, preserving the preceding CodeGemma correction. Only the extracted comment columns for CodeLlama, CodeQwen 1.5, and GraniteCode change. Fourteen comments are truncated before their first generation-control token: six CodeLlama <EOT> endings, two CodeLlama <SUF> continuations, five GraniteCode <|endoftext|> endings, and one CodeQwen <|endoftext|> ending. The preceding text and whitespace are preserved exactly.

For Greek file 777248_1, CodeQwen's extracted-comment field incorrectly equals its entire raw generation. Its exact stored masked_data prompt prefix (791 characters, ending in <fim_middle>) is removed, preserving the 1,378-character generated completion exactly.

LanguageCodeLlamaCodeQwen 1.5GraniteCode
English200
Greek625

These 15 corrections introduce no new blank completions. Raw generations, prompts, references, expert labels, error codes, row order, CodeGemma, and StarCoder2 values are unchanged. Previously computed neural scores for changed comments must be recomputed.