Team Ai
Datasetpublic

violetakastreva/line-level-code-vs-text-classification

This dataset was created for SemEval-2026 Task 13, which focuses on distinguishing machine-generated code from human-written code across multiple programming languages and domains. While the original SemEval task operates at the code snippet level, this dataset provides line-level annotations that enable finer-grained analysis of how code-like and text-like content is distributed within mixed inputs. The dataset is intended to support research in machine-generated code detection, robust… See the full description on the dataset page: https://huggingface.co/datasets/violetakastreva/line-level-code-vs-text-classification.

sourceHugging Faceodblupdated 8mo agoView on Hugging Face
0likes12downloads
3 commits on main
74cca878mo ago

Update README.md

violetakastreva
2a99e0a8mo ago

Upload code-vs-text.csv

violetakastreva
30cd7638mo ago

initial commit

violetakastreva