AISE-TUDelft/multilingual-code-comments-fixed-8
Fixed-8 Based on fixed-7 revision 14e85fe00a8b284cd226c58281ddd8e6b990b190. Replaces six Greek rows with missing expert labels with six newly labelled samples. All five language configurations retain 500 training rows (2,500 total). All other rows are unchanged. Removed ID Replacement ID 8000_5 1056_0 8000_14 4357_8 8000_15 4848_9 8000_16 29069_13 8000_17 1385_4 8000_18 5142_0 All 500 Greek rows now have all five expert accuracy labels. Original… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/multilingual-code-comments-fixed-8.
Fixed-8
Based on fixed-7 revision 14e85fe00a8b284cd226c58281ddd8e6b990b190. Replaces six Greek rows with missing expert labels with six newly labelled samples. All five language configurations retain 500 training rows (2,500 total). All other rows are unchanged.
All 500 Greek rows now have all five expert accuracy labels. Original supplied predictions and error-code strings are preserved; blank error-code fields are retained as provided. This release does not remove pre-existing file reuse or certify absence of partial code clones. Scores and judge results from fixed-7 must be recomputed for the six replacement IDs.
CodeGemma comment extraction correction
Corrected from dataset revision 292880b74aa2ec58dc1b00524a70b060359bf0fa. Only predicted_comment_google/codegemma-7b was changed: each affected completion is truncated immediately before its first <|file_separator|> token. The token and all subsequent generated text are removed; the preceding text and whitespace are preserved exactly. Raw predict_google/codegemma-7b generations, prompts, reference comments, expert labels, error codes, row order, and every other generator's columns are unchanged.
11 affected completions contain only whitespace or no text before the boundary; these are preserved as such, with their supplied expert labels unchanged. Previously computed neural scores for changed comments must be recomputed from these corrected inputs.
Remaining comment extraction corrections
Corrected from revision d02cf177485908c0d26e71c26f4a584c0e38dd5d, preserving the preceding CodeGemma correction. Only the extracted comment columns for CodeLlama, CodeQwen 1.5, and GraniteCode change. Fourteen comments are truncated before their first generation-control token: six CodeLlama <EOT> endings, two CodeLlama <SUF> continuations, five GraniteCode <|endoftext|> endings, and one CodeQwen <|endoftext|> ending. The preceding text and whitespace are preserved exactly.
For Greek file 777248_1, CodeQwen's extracted-comment field incorrectly equals its entire raw generation. Its exact stored masked_data prompt prefix (791 characters, ending in <fim_middle>) is removed, preserving the 1,378-character generated completion exactly.
These 15 corrections introduce no new blank completions. Raw generations, prompts, references, expert labels, error codes, row order, CodeGemma, and StarCoder2 values are unchanged. Previously computed neural scores for changed comments must be recomputed.
