Team Ai
Datasetpublic

Silasimo/GTSinger-EN-Annotations

These customized annotation files are based on, and curated for, the English partition of the GTSinger dataset. The annotation files were created for singing-oriented forced alignment experiments in our SynthGT project. The original GTSinger annotations have been processed by removing stress markers, lowerchasing phonemes, and substituting all instances of AP with SP, to share the same vocabulary as our SynthGT dataset. The annotations are distributed in train/valid/test splits as .csv… See the full description on the dataset page: https://huggingface.co/datasets/Silasimo/GTSinger-EN-Annotations.

sourceHugging Facecc-by-nc-sa-4.0updated 2d agoView on Hugging Face
0likes28downloads
Dataset Card

![Hugging Face](https://huggingface.co/datasets/AaronZ345/GTSinger) ![Hugging Face](https://huggingface.co/datasets/Silasimo/SynthGT) ![Hugging Face](https://huggingface.co/collections/Silasimo/synthgt-forced-alignment-experiments) ![GitHub](https://github.com/SilasAntonisen/SynthGT)

These customized annotation files are based on, and curated for, the English partition of the GTSinger dataset. The annotation files were created for singing-oriented forced alignment experiments in our SynthGT project. The original GTSinger annotations have been processed by removing stress markers, lowerchasing phonemes, and substituting all instances of AP with SP, to share the same vocabulary as our SynthGT dataset. The annotations are distributed in train/valid/test splits as .csv files, which include the associated audio file name, the phoneme sequence, and the duration of each phoneme in the sequence. Additionally, the test split includes the annotations in .TextGrid format, and .lab files with only the phoneme sequence and not the phoneme durations, which is convenient for model evaluation.

Forced alignment models which we have trained with these annotation files can be found in our model collection.

Citation

If you use these files in your own experiments, please cite out work:

bibtex
@unpublished{Antonisen2026SynthGT,
  author = {Silas Antonisen and Iván López-Espejo},
  title  = {Synthesizing Ground Truth Phoneme Boundaries with SynthGT: A Synthetically Generated Solo-Singing Corpus},
  note   = {Submitted to IEEE Transactions on Audio, Speech and Language Processing},
  year   = {2026}
}

As well as the original GTSinger authors:

bibtex
@inproceedings{GTSinger,
  title = {{GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks}},
  booktitle = {Proceedings of the 38th Conference on Neural Information Processing Systems (NeurIPS)},
  author = {Yu Zhang and Changhao Pan and Wenxiang Guo and Ruiqi Li and Zhiyuan Zhu and Jialei Wang and Wenhao Xu and Jingyu Lu and Zhiqing Hong and Chuxin Wang and LiChao Zhang and Jinzheng He and Ziyue Jiang and Yuxin Chen and Chen Yang and Jiecheng Zhou and Xinyu Cheng and Zhou Zhao},
  year = {2024},
  month = {December 10--15},
  pages = {1--24},
  address = {Vancouver, Canada}
}