abdo1819/arabic-english-code-switching-review-annotations
Review Annotations for Arabic-English Code-Switching Speech This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts. The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index. Coverage and outcomes The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.
Review Annotations for Arabic-English Code-Switching Speech
This metadata-only dataset publishes review decisions and transcript-correction deltas for `MohamedRashad/arabic-english-code-switching`. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts.
The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index.
Coverage and outcomes
The upstream revision has 12,480 rows. The local review workflow imported 12,473 rows; upstream indices 116, 869, 6158, 10304, 10487, 11561, and 11618 were not imported.
The accepted set contains 8,988 rows: 7,264 train, 862 validation, and 862 test. Only accepted transcripts that differ from the pinned upstream transcript include corrected_text; all other rows use a null correction delta.
Loading and joining to upstream audio
from datasets import load_dataset
annotations = load_dataset(
"abdo1819/arabic-english-code-switching-review-annotations",
split="train",
)
upstream = load_dataset(
"MohamedRashad/arabic-english-code-switching",
split="train",
revision="4a3bffc45219c35949470de32b8d4cb328b0ce11",
)
annotation_by_row = {row["upstream_row_index"]: row for row in annotations}
def accepted_rows():
for row_index, source_row in enumerate(upstream):
annotation = annotation_by_row.get(row_index)
if not annotation or annotation["accepted"] is not True:
continue
yield {
"audio": source_row["audio"],
"text": annotation["corrected_text"] or source_row["sentence"],
"split": annotation["canonical_split"],
"review_status": annotation["review_status"],
}This join retrieves audio from the upstream repository at use time; the annotation repository does not redistribute it.
Fields
id: stable annotation IDupstream_dataset,upstream_revision,upstream_config,upstream_split: pinned parent identityupstream_row_indexandupstream_row_url: join key and human-readable source linklocal_review_source_index: ordering in the 12,473-row local review inputsource_text_sha256andsource_text_length: source-transcript verification fields without republishing the textreview_status: detailed decision statusreviewed: whether an explicit decision existsaccepted:true,false, or null for unreviewed rowsreview_method:manual,gemini_assisted,rule_based, orunreviewedrejection_reason: normalized rejection categorycanonical_split:train,validation,test, or empty for non-accepted rowstext_modified: whether the accepted review transcript differs exactly from the pinned source transcriptcorrected_text: correction delta only; null when the source transcript is retainedreviewed_text_sha256: checksum of the final accepted transcript- automatic-review audit fields: rule, threshold, selected variant, direct/structured normalized WER, and duplicate parent row
decision_updated_at: review decision timestamp
Review process
Rows enter the accepted set only through keep or gemini_keep decisions. Manual decisions can correct transcript text. Automated assists include:
- duplicate rejection based on canonical audio identity or normalized transcript;
- rejection of Arabic-only or English-only transcripts for the code-switching target;
- Gemini-assisted keep when either no-thinking transcription variant has normalized WER at most
0.2; - Gemini high-WER rejection when both no-thinking variants exceed normalized WER
1.5.
Machine-assisted keeps are not equivalent to manual verification. The project’s manual-only audit found high precision but limited recall at the 0.2 auto-keep threshold.
Licensing and attribution
The upstream dataset currently declares gpl without specifying a version. This annotation repository does not redistribute upstream audio and minimizes transcript reproduction by publishing only correction deltas. The upstream dataset’s terms still apply when users retrieve or use its audio and transcripts.
Abdelrahman R. Hashem releases the independently authored review labels, audit metadata, and correction deltas under CC BY 4.0. This license does not grant rights to upstream audio or unchanged upstream transcripts.
This section records provenance and risks; it is not legal advice.
Limitations
- 2,376 imported rows remain unreviewed.
- Seven upstream rows were not present in the local review input.
- Automatic accept/reject rules can make mistakes and should remain distinguishable from manual decisions.
- Reliable speaker identifiers are unavailable, so canonical splits cannot guarantee speaker disjointness.
- Correction deltas can contain names, brands, URLs, or other entities already present or spoken in the source material.
- Row-index joins require the pinned upstream revision; later upstream revisions may reorder or change rows.
Citation
@dataset{hashem2026codeswitchreview,
author = {Hashem, Abdelrahman R.},
title = {Review Annotations for Arabic-English Code-Switching Speech},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations}
}Users must also cite the upstream dataset.
Contact
Abdelrahman R. Hashem — arh13@fayoum.edu.eg
