Team Ai
Datasetpublic

RareConcepts/suno-reggae-test-dataset

Suno Patois Reggae Test Set 299 patois-language reggae and dancehall tracks with style captions and structured lyrics, laid out for SimpleTuner's textfile audio caption strategy. Built as a small, high-consistency probe set for text-to-audio training runs — not a general-purpose music corpus. Rights and provenance Every track here was generated by a third-party Suno user, and rights in the audio and lyrics remain with those creators. Nothing in this repository is… See the full description on the dataset page: https://huggingface.co/datasets/RareConcepts/suno-reggae-test-dataset.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes396downloads
Dataset Card

Suno Patois Reggae Test Set

299 patois-language reggae and dancehall tracks with style captions and structured lyrics, laid out for SimpleTuner's textfile audio caption strategy. Built as a small, high-consistency probe set for text-to-audio training runs — not a general-purpose music corpus.

Rights and provenance

Every track here was generated by a third-party Suno user, and rights in the audio and lyrics remain with those creators. Nothing in this repository is owned by the uploader, and no license is granted over the underlying content. The license: other tag above reflects that the material is redistributed without a rights grant, not that it is freely reusable.

  • —299 tracks from 213 distinct creators.
  • —manifest.json carries per-track attribution: the creator handle, the Suno clip id, and a source_url linking back to the original song page.
  • —Lyrics are human-authored by those creators. Audio is model-generated by Suno.

If you are one of the creators represented here and want your work removed, open a discussion on this repository and it will be taken down.

Layout

Flat directory of matched triples sharing a basename:

<name>.mp3      audio
<name>.txt      style caption  (SimpleTuner caption_strategy: textfile)
<name>.lyrics   structured lyrics
manifest.json   per-track metadata and attribution

Basenames are <title_slug>_<first 8 chars of clip id>. The id suffix is load bearing — Suno titles collide frequently, and without it a collision would pair one track's audio against another's lyrics.

.txt — style caption

A comma-delimited narrative in the style of SimpleTuner's examples/minimaxmusic-prompts.json, derived from the creator's own style tags plus grounded additions (patois vocal delivery, reggae feel, a vocal clause, repeated chorus hook where the lyrics contain one).

No BPM or musical key is synthesized. Suno's metadata has no tempo field, so a caption states tempo only where the creator wrote it themselves. Captions here are never fabricated beyond what the source metadata supports.

.lyrics — structured lyrics

Section markers on their own lines, newlines preserved:

[verse]
...
[chorus]
...

Headers are normalized to lowercase canonical forms ([Verse 1] → [verse], [Pre-Chorus] → [pre-chorus]). Lyric text is otherwise untouched — patois orthography is preserved exactly as the creator wrote it and has not been normalized toward standard English.

All 299 files carry markers authored by the original creators. No section structure in this set is machine-derived.

Selection criteria

Sourced via Suno's public tag_song search across the terms patois, dancehall, roots reggae, ragga, lovers rock, rocksteady, dub, and reggae, then filtered to require all of:

filterrationale
patois in the creator's style tagsthe target concept
a reggae-family tagexcludes patois vocals over unrelated genres
non-empty lyricsinstrumentals defeat the purpose of a vocal set
lyrics contain [section] markersstructure must be authored, not inferred
lyrics ≥ 200 charsdrops degenerate stubs
style tags ≥ 40 charsfloors caption richness
unique style-tag stringone creator reusing a prompt across a release would otherwise contribute many identical captions

Survivors were ranked by upvote count and the top 300 taken. One track's audio returned HTTP 403 during fetch and was dropped, giving 299.

Statistics

tracks299
total audio16.8 hours
durationmedian 192s (38–422s)
creators213
caption lengthmedian 354 chars (45–1000)
lyrics lengthmedian 1835 chars (276–5000)
duplicate captions0
tracks missing any of the three files0

Intended use and limitations

Built to test whether a text-to-audio model can pick up a specific vocal register and genre from a small set. It is not a benchmark and not balanced for evaluation.

  • —Small and narrow. 299 tracks in one genre and register.
  • —Synthetic audio. All audio is Suno model output, so it carries that model's artifacts and production signature rather than recorded music.
  • —Captions inherit creator vocabulary. Tag conventions vary between creators; some captions are terse, others detailed.
  • —Explicit content. Dancehall has an explicit lyrical tradition and this set reflects it. Filter on lyrics before use if that matters for your run.
  • —Popularity-biased. Selection by upvote count skews toward what performs on Suno, not toward a representative sample of the genre.

SimpleTuner configuration

json
{
  "id": "patois-reggae",
  "type": "local",
  "dataset_type": "audio",
  "instance_data_dir": "/path/to/suno-reggae-test-dataset",
  "caption_strategy": "textfile",
  "audio": { "lyrics_filename_format": "{filename}.lyrics" },
  "cache_dir_vae": "cache/vae/{model_family}/patois-reggae"
}

lyrics_filename_format must be set explicitly for any model family other than ACE-Step. SimpleTuner only applies the {filename}.lyrics default when model_family == "ace_step", so on other families the lyrics silently fail to load without it.