Team Ai
Datasetpublic

violetakastreva/line-level-code-vs-text-classification

This dataset was created for SemEval-2026 Task 13, which focuses on distinguishing machine-generated code from human-written code across multiple programming languages and domains. While the original SemEval task operates at the code snippet level, this dataset provides line-level annotations that enable finer-grained analysis of how code-like and text-like content is distributed within mixed inputs. The dataset is intended to support research in machine-generated code detection, robust… See the full description on the dataset page: https://huggingface.co/datasets/violetakastreva/line-level-code-vs-text-classification.

sourceHugging Faceodblupdated 8mo agoView on Hugging Face
0likes12downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
violetakastreva/line-level-code-vs-text-classification · Team Ai