datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smart_contract_code_commentsmultilingual-code-comments-fixed-8
Multilingual code comments
This dataset contains 500 source-code examples for each of Chinese, Dutch, English, Greek and Polish (2,500 examples total), with comments generated by five models and human correctness ratings and error annotations. Each language has a train split.
Human ratings use Correct, Partial and Incorrect. Error fields contain comma-separated taxonomy codes. A generated comment can carry multiple error codes.
Annotation interpretation
Stored… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/multilingual-code-comments-fixed-8.multilingual-code-comments
A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics
This dataset helps us understand how Large Language Models (LLMs) can create code comments in different languages. While LLMs are good at coding tasks in English, we don't know much about how well they work in other languages. This dataset, along with our research, studies how LLMs generate code comments in English, Chinese, Dutch, Polish, and Greek. In our case, we have… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/multilingual-code-comments.multilingual-code-comments-fixed-5multilingual-code-comments-fixed-6multilingual-code-comments-fixedmultilingual-code-comments-fixed-2multilingual-code-comments-fixed-3multilingual-code-comments-fixed-4multilingual-code-comments-fixed-7smart_contract_code_commentscode-comment-syntheticmultilingual-code-comments_English_Correct
