AISE-TUDelft/leading-comments
Leading Comments We release Leading Comments, a collection of opening comment blocks from source code in The Stack v1, GitHub-Code, CodeParrot, The Pile, and RedPajama. The dataset brings together file-level license notices, copyright statements, redistribution conditions, and other introductory comments to support research on code licensing and dataset curation. Paper · Preprint · Code and replication package Getting started Install the Hugging Face Datasets… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/leading-comments.
Leading Comments
We release Leading Comments, a collection of opening comment blocks from source code in The Stack v1, GitHub-Code, CodeParrot, The Pile, and RedPajama. The dataset brings together file-level license notices, copyright statements, redistribution conditions, and other introductory comments to support research on code licensing and dataset curation.
Paper · Preprint · Code and replication package
Getting started
Install the Hugging Face Datasets library:
pip install datasetsLoad a subset by its configuration name. Use streaming to start exploring without downloading the full subset:
from datasets import load_dataset
comments = load_dataset(
"AISE-TUDelft/leading-comments",
"TheStack",
split="train",
streaming=True,
)
for example in comments.take(5):
print(example["comments"])Omit streaming=True to download the subset for local processing.
Data
The comments are organized by source dataset and stored in Parquet format. Each configuration provides a train split with a single column:
How we collected the comments
We extract the first comment block when it begins within the first 20 characters of a source file, retaining the block beyond that initial window. This captures introductory text such as license headers, copyright notices, file descriptions, and instructions about sharing or using the code.
We use StarPII to identify and redact personally identifiable information before release. The extraction and processing resources are available in our replication package.
Using the dataset
Leading Comments supports the development and evaluation of methods for detecting license notices, recognizing redistribution conditions, and studying the information developers place at the beginning of source files. It can also be used to explore how file-level notices complement repository-level metadata when curating code datasets.
Citation
If you use Leading Comments, please cite our FORGE 2024 paper, An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets:
@inproceedings{katzy2024exploratory,
title = {An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets},
author = {Katzy, Jonathan and Popescu, Răzvan-Mihai and van Deursen, Arie and Izadi, Maliheh},
booktitle = {Proceedings of the 2024 IEEE/ACM First International Conference on AI Foundation Models and Software Engineering},
series = {FORGE '24},
pages = {74--85},
year = {2024},
publisher = {Association for Computing Machinery},
doi = {10.1145/3650105.3652298},
url = {https://doi.org/10.1145/3650105.3652298}
}Authors
Jonathan Katzy, Răzvan-Mihai Popescu, Arie van Deursen, and Maliheh Izadi — Delft University of Technology.
For questions, use the dataset discussions.
