QUD-Technologies/quran-alignment-benchmark
Quran Recitation Alignment Benchmark Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here. 16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim ·… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quran-alignment-benchmark.
Quran Recitation Alignment Benchmark
 
Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here.
16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim · murattal, mujawwad, hadr and muallim · studio, prayer and teaching captures.
Using the dataset
One config per corpus version, one split.
from datasets import load_dataset
ds = load_dataset("QUD-Technologies/quran-alignment-benchmark", "v1", split="test")
row = ds[0]
row["audio"] # decoded audio; use Audio(decode=False) to get the MP3 bytes instead
row["truth"]["words"] # [{"word": "84:1:1", "start_s": 5.78, "end_s": 6.25}, ...]
row["segments"] # [{"start_s", "end_s", "first_word", "last_word"}, ...]To run an aligner on the recordings and score it, use the benchmark package rather than reading the parquet yourself:
pip install "quran-alignment-benchmark[corpus]"
qab fetch --corpus v1 --out corpus/ # <id>.mp3 + <id>.json per recording
qab score submissions/ --corpus v1 # the same report the leaderboard showsColumns
Times are seconds from the start of the audio. A repeated word appears once per recitation. Full column definitions, annotation method and known gaps: docs/CORPUS.md.
Versioning
A config is immutable once a score has been published against it. A re-annotation or an added recording becomes the next config (v2). Scores always name the config they were computed on.
License
CC BY 4.0 for the annotations (truth, segments) and descriptive columns, and for the benchmark repository. The audio recordings are not covered by this license.
Corpus issues and new audio
Report an audio or ground-truth problem through a GitHub issue.
To suggest new audio, provide only:
- An audio link or file.
- The reciter's name.
- Why it would be a useful addition to the corpus.
No other metadata or annotations are needed. If accepted, the recording will be ingested, annotated, and released in the next corpus version.
