Team Ai
Datasetpublic

AISE-TUDelft/leading-comments

Leading Comments We release Leading Comments, a collection of opening comment blocks from source code in The Stack v1, GitHub-Code, CodeParrot, The Pile, and RedPajama. The dataset brings together file-level license notices, copyright statements, redistribution conditions, and other introductory comments to support research on code licensing and dataset curation. Paper · Preprint · Code and replication package Getting started Install the Hugging Face Datasets… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/leading-comments.

sourceHugging Faceupdated 9d agoView on Hugging Face
0likes1.1kdownloads
Dataset Card

Leading Comments

We release Leading Comments, a collection of opening comment blocks from source code in The Stack v1, GitHub-Code, CodeParrot, The Pile, and RedPajama. The dataset brings together file-level license notices, copyright statements, redistribution conditions, and other introductory comments to support research on code licensing and dataset curation.

Paper · Preprint · Code and replication package

Getting started

Install the Hugging Face Datasets library:

bash
pip install datasets

Load a subset by its configuration name. Use streaming to start exploring without downloading the full subset:

python
from datasets import load_dataset

comments = load_dataset(
    "AISE-TUDelft/leading-comments",
    "TheStack",
    split="train",
    streaming=True,
)

for example in comments.take(5):
    print(example["comments"])

Omit streaming=True to download the subset for local processing.

Data

The comments are organized by source dataset and stored in Parquet format. Each configuration provides a train split with a single column:

ColumnTypeContents
commentsstringThe opening comment block extracted from a source file.
Source datasetConfiguration name(s)Rows per configuration
The Stack v1TheStack77,595,559
GitHub-CodeGitHubCode45,301,797
CodeParrotCodeParrot, CodeParrotComments14,372,397
The PileThePile, ThePileComments6,794,995
RedPajamaRedPajama, RedPajamaComments2,281,378

How we collected the comments

We extract the first comment block when it begins within the first 20 characters of a source file, retaining the block beyond that initial window. This captures introductory text such as license headers, copyright notices, file descriptions, and instructions about sharing or using the code.

We use StarPII to identify and redact personally identifiable information before release. The extraction and processing resources are available in our replication package.

Using the dataset

Leading Comments supports the development and evaluation of methods for detecting license notices, recognizing redistribution conditions, and studying the information developers place at the beginning of source files. It can also be used to explore how file-level notices complement repository-level metadata when curating code datasets.

Citation

If you use Leading Comments, please cite our FORGE 2024 paper, An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets:

bibtex
@inproceedings{katzy2024exploratory,
  title = {An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets},
  author = {Katzy, Jonathan and Popescu, Răzvan-Mihai and van Deursen, Arie and Izadi, Maliheh},
  booktitle = {Proceedings of the 2024 IEEE/ACM First International Conference on AI Foundation Models and Software Engineering},
  series = {FORGE '24},
  pages = {74--85},
  year = {2024},
  publisher = {Association for Computing Machinery},
  doi = {10.1145/3650105.3652298},
  url = {https://doi.org/10.1145/3650105.3652298}
}

Authors

Jonathan Katzy, Răzvan-Mihai Popescu, Arie van Deursen, and Maliheh Izadi — Delft University of Technology.

For questions, use the dataset discussions.