Team Ai
Datasetpublic

years0/multilingual-code-switching-bench

Multilingual Code-Switching & Dialectal Evaluation Benchmark (Fatima Fellowship Application) Question 1: Critical Blind Spot & Capability Gap Standard NLP benchmarks (MMLU, GSM8K, HumanEval) evaluate language models on clean, standardized, monolingual inputs. However, for billions of global speakers, everyday digital communication occurs in low-resource code-switched vernaculars (e.g., Franco-Arabic/Arabizi, Singlish, Taglish, Hinglish, Naija Pidgin, Sheng… See the full description on the dataset page: https://huggingface.co/datasets/years0/multilingual-code-switching-bench.

sourceHugging Facemitupdated 19h agoView on Hugging Face
0likes23downloads
Dataset Card

Multilingual Code-Switching & Dialectal Evaluation Benchmark (Fatima Fellowship Application)

Question 1: Critical Blind Spot & Capability Gap

Standard NLP benchmarks (MMLU, GSM8K, HumanEval) evaluate language models on clean, standardized, monolingual inputs. However, for billions of global speakers, everyday digital communication occurs in low-resource code-switched vernaculars (e.g., Franco-Arabic/Arabizi, Singlish, Taglish, Hinglish, Naija Pidgin, Sheng, Spanglish, and Papiamento).

Current frontier small language models (0.6B–6B parameters) exhibit three systematic failure modes in these linguistic regimes:

  1. 1.Semantic Hallucination & False Rules: When encountering regional idioms or transliterated vernaculars, models hallucinate fake explanations to justify literal translations (e.g., asserting that "pregnant" is a common Spanish slang term for "anxious").
  2. 2.Subword Fragmentation & Transliteration Penalty: Tokenizers optimized for standard Western scripts fragment Romanized or creole dialects into excessive subwords, degrading downstream reasoning.
  3. 3.Safety Guardrail & Tone Misalignment: Informal dialectal praise, regional exclamations, or colloquial slang are frequently misclassified or over-sanitized into generic boilerplate.

Question 2: Systematic Evaluation Narrative (google/gemma-2-2b-it)

Using a custom 10-item multilingual test suite covering 8 global dialect systems, we evaluated google/gemma-2-2b-it (2.6B parameters) in 4-bit precision.

Selected Empirical Results

IDDialect / Language PairPromptExpected IntentModel Output / BehaviorFailure Category
1Arabizi + EnglishYa zalameh, wallahi el mashroo3 dah 3al-alla, bas rabna hayastor.Project uncertainty / reliance on God"I swear it's been a long time, but I'm back."Hallucination: Completely missed project context; invented a return narrative.
3SinglishDon't keyi like that lah, the boss already kancheong, later he pok kai then how?Warning against reckless action causing bankruptcy/ruin.Mapped kancheong to "cold shoulder" and pok kai to "scolding".Pragmatic Failure: Failed to recognize financial ruin vs scolding.
5TaglishSuper nakakagigil yung deployment test... pero keribels lang kasi resolved na agad.Frustrating test, but manageable (keribels).Defined keribels as meaning "stressful".Polarity Inversion: Direct opposite meaning of slang term.
9SpanglishEstoy super embarazada por el error que cometí...Embarrassment over a meeting blunder.Claimed "super pregnant" is Spanish slang for "extreme anxiety".False Cognate Rationalization: Invented a false linguistic rule to defend literal error.

Question 3: Path Forward & Proposed Solutions

To bridge the performance gap in sub-6B open-weights models without increasing parameters:

  1. 1.Vocabulary Expansion & Byte-Level Subword Merging: Retrain base tokenizers on multi-dialectal social media corpora (X, Telegram, Reddit) to merge frequent Romanized and creole n-grams, reducing fragmentation penalties.
  2. 2.Participatory Community Data Curation: Shift away from synthetic LLM-generated dialect corpora (which reproduce alignment biases). Recruit native speakers across dialect communities to curate authentic, pragmatically annotated conversational pairs.
  3. 3.Direct Preference Optimization (DPO) on Pragmatic Alignment: Fine-tune models using LoRA adapters with preference pairs (y_win, y_lose), where y_lose represents literal translations or invented explanations, and y_win represents pragmatically accurate, context-aware responses.

Repository Artifacts

  • —multilingual_evaluation_results.json: Raw evaluation logs and prompt outputs.
  • —evaluation_notebook.ipynb: Runnable Google Colab notebook used for benchmark execution.