Team Ai
Datasetpublic

MehdiAstaraki/progressive-code-switch

MehdiAstaraki/progressive-code-switch Progressive code-switching retrieval-decay benchmark. Each base patent yields a cumulative ladder of documents (base__r0 clean → base__rN), where each step swaps one more chemistry term into another language / spelling / ChEBI form. One fixed question per base (about the step-1 term) is reused for every depth; the qrels carry the depth so you can measure how retrieval decays as more terms are code-switched. Configs: corpus (ladder variant… See the full description on the dataset page: https://huggingface.co/datasets/MehdiAstaraki/progressive-code-switch.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes6downloads
Dataset Card

MehdiAstaraki/progressive-code-switch

Progressive code-switching retrieval-decay benchmark. Each base patent yields a cumulative ladder of documents (base__r0 clean → base__rN), where each step swaps one more chemistry term into another language / spelling / ChEBI form. One fixed question per base (about the step-1 term) is reused for every depth; the qrels carry the depth so you can measure how retrieval decays as more terms are code-switched.

Configs: corpus (ladder variant documents, with depth + mode_added), queries (one per base), qrels (query→variant, score 1, with depth). The retrieval haystack is this corpus PLUS a shared patent corpus passed at eval time.