Team Ai
Datasetpublic

abdo1819/arabic-english-code-switching-review-annotations

Review Annotations for Arabic-English Code-Switching Speech This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts. The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index. Coverage and outcomes The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes26downloads
Dataset Card

Review Annotations for Arabic-English Code-Switching Speech

This metadata-only dataset publishes review decisions and transcript-correction deltas for `MohamedRashad/arabic-english-code-switching`. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts.

The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index.

Coverage and outcomes

The upstream revision has 12,480 rows. The local review workflow imported 12,473 rows; upstream indices 116, 869, 6158, 10304, 10487, 11561, and 11618 were not imported.

OutcomeRows
Manual keep3,487
Gemini-assisted keep5,501
Manual remove328
Rule-based duplicate rejection81
Rule-based no-switching rejection565
Gemini high-WER rejection135
Unreviewed2,376

The accepted set contains 8,988 rows: 7,264 train, 862 validation, and 862 test. Only accepted transcripts that differ from the pinned upstream transcript include corrected_text; all other rows use a null correction delta.

Loading and joining to upstream audio

python
from datasets import load_dataset

annotations = load_dataset(
    "abdo1819/arabic-english-code-switching-review-annotations",
    split="train",
)

upstream = load_dataset(
    "MohamedRashad/arabic-english-code-switching",
    split="train",
    revision="4a3bffc45219c35949470de32b8d4cb328b0ce11",
)

annotation_by_row = {row["upstream_row_index"]: row for row in annotations}

def accepted_rows():
    for row_index, source_row in enumerate(upstream):
        annotation = annotation_by_row.get(row_index)
        if not annotation or annotation["accepted"] is not True:
            continue
        yield {
            "audio": source_row["audio"],
            "text": annotation["corrected_text"] or source_row["sentence"],
            "split": annotation["canonical_split"],
            "review_status": annotation["review_status"],
        }

This join retrieves audio from the upstream repository at use time; the annotation repository does not redistribute it.

Fields

  • —id: stable annotation ID
  • —upstream_dataset, upstream_revision, upstream_config, upstream_split: pinned parent identity
  • —upstream_row_index and upstream_row_url: join key and human-readable source link
  • —local_review_source_index: ordering in the 12,473-row local review input
  • —source_text_sha256 and source_text_length: source-transcript verification fields without republishing the text
  • —review_status: detailed decision status
  • —reviewed: whether an explicit decision exists
  • —accepted: true, false, or null for unreviewed rows
  • —review_method: manual, gemini_assisted, rule_based, or unreviewed
  • —rejection_reason: normalized rejection category
  • —canonical_split: train, validation, test, or empty for non-accepted rows
  • —text_modified: whether the accepted review transcript differs exactly from the pinned source transcript
  • —corrected_text: correction delta only; null when the source transcript is retained
  • —reviewed_text_sha256: checksum of the final accepted transcript
  • —automatic-review audit fields: rule, threshold, selected variant, direct/structured normalized WER, and duplicate parent row
  • —decision_updated_at: review decision timestamp

Review process

Rows enter the accepted set only through keep or gemini_keep decisions. Manual decisions can correct transcript text. Automated assists include:

  • —duplicate rejection based on canonical audio identity or normalized transcript;
  • —rejection of Arabic-only or English-only transcripts for the code-switching target;
  • —Gemini-assisted keep when either no-thinking transcription variant has normalized WER at most 0.2;
  • —Gemini high-WER rejection when both no-thinking variants exceed normalized WER 1.5.

Machine-assisted keeps are not equivalent to manual verification. The project’s manual-only audit found high precision but limited recall at the 0.2 auto-keep threshold.

Licensing and attribution

The upstream dataset currently declares gpl without specifying a version. This annotation repository does not redistribute upstream audio and minimizes transcript reproduction by publishing only correction deltas. The upstream dataset’s terms still apply when users retrieve or use its audio and transcripts.

Abdelrahman R. Hashem releases the independently authored review labels, audit metadata, and correction deltas under CC BY 4.0. This license does not grant rights to upstream audio or unchanged upstream transcripts.

This section records provenance and risks; it is not legal advice.

Limitations

  • —2,376 imported rows remain unreviewed.
  • —Seven upstream rows were not present in the local review input.
  • —Automatic accept/reject rules can make mistakes and should remain distinguishable from manual decisions.
  • —Reliable speaker identifiers are unavailable, so canonical splits cannot guarantee speaker disjointness.
  • —Correction deltas can contain names, brands, URLs, or other entities already present or spoken in the source material.
  • —Row-index joins require the pinned upstream revision; later upstream revisions may reorder or change rows.

Citation

bibtex
@dataset{hashem2026codeswitchreview,
  author    = {Hashem, Abdelrahman R.},
  title     = {Review Annotations for Arabic-English Code-Switching Speech},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations}
}

Users must also cite the upstream dataset.

Contact

Abdelrahman R. Hashem — arh13@fayoum.edu.eg