surrey-nlp/dialect-preferences
DiaLLM — Pooled Preference Dataset (Implicit Thread) Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 45,690 preference pairs, pooling all three variety-specific sets (Australian, Northern British, Indian) without variety targeting. Used for implicit-thread DPO training, where the three varieties are pooled rather than targeted individually, preserving the variety-agnostic objective of that thread.… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/dialect-preferences.
DiaLLM — Pooled Preference Dataset (Implicit Thread)
Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main).

45,690 preference pairs, pooling all three variety-specific sets (Australian, Northern British, Indian) without variety targeting. Used for implicit-thread DPO training, where the three varieties are pooled rather than targeted individually, preserving the variety-agnostic objective of that thread.
Construction
Built from the UltraFeedback preference dataset (Cui et al., 2023), via Argilla's cleaned/binarized release (`argilla/ultrafeedback-binarized-preferences-cleaned`, MIT licensed): the originally-preferred completion is transformed into a dialectal variant using Multi-VALUE (Ziems et al., 2023), based on eWAVE morphosyntactic features. Code blocks are preserved verbatim during conversion.
Columns
Code, checkpoints, linguistic-analysis toolkit: https://github.com/surrey-nlp/diallm
Paper: https://arxiv.org/abs/2607.07669
Citation
@article{painter2026diallm,
title = {DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation},
author = {Painter, Jordan and Srirag, Dipankar and Kappiyath, Adarsh and Kanojia, Diptesh and Joshi, Aditya and Yin, Lu},
year = {2026},
eprint = {2607.07669},
archivePrefix = {arXiv}
}