Team Ai
Datasetpublic

mikezhu/chord-experiments-data

CHORD experiment corpora Text corpora and cached encoder features behind every table and figure of the paper Coherence-Aware Distributional Evaluation of Open-Ended Text Generation: the counterfactual evaluation set, the unconditional-generation samples and human reference pools, and the prefix-continuation, human-agreement, QA-faithfulness and appendix texts. The tree mirrors the experiments repository (CHORD-Experiment), so after python scripts/download_data.py --repo… See the full description on the dataset page: https://huggingface.co/datasets/mikezhu/chord-experiments-data.

sourceHugging Faceotherupdated 11d agoView on Hugging Face
0likes190downloads
Dataset Card

CHORD experiment corpora

Text corpora and cached encoder features behind every table and figure of the paper Coherence-Aware Distributional Evaluation of Open-Ended Text Generation: the counterfactual evaluation set, the unconditional-generation samples and human reference pools, and the prefix-continuation, human-agreement, QA-faithfulness and appendix texts. The tree mirrors the experiments repository (CHORD-Experiment), so after

bash
python scripts/download_data.py --repo mikezhu/chord-experiments-data            # texts
python scripts/download_data.py --repo mikezhu/chord-experiments-data --features # + cached features

every config resolves. MANIFEST-experiments.tsv lists every file with its size and sha256. The students' training data is in a separate dataset (chord-distill-data).

Code: https://github.com/MAPS-research/CHORD

Contents

PathWhat
outputs/counterfactual/meta_eval/counterfactual evaluation set (Table 1): texts per condition and the conditions manifest
outputs/casestudy/unconditional_generation/Table 2: 10 seeds x 500 samples per generator, human reference / held-out pools
outputs/casestudy/single_fold/single-fold unconditional corpora (quick student check)
outputs/casestudy/prefix_continuation/prefixes, human continuations and generator continuations (Fig. 5)
outputs/casestudy/human_agreement/GPT-2 texts with public human pairwise judgments and their Bradley-Terry scores (Table 3)
outputs/casestudy/qa_faithfulness/source-conditioned QA-faithfulness texts (appendix)
outputs/experiments/position_robustness/failure-position corpus (appendix)
outputs/casestudy/unconditional_generation/features/, cache/ (features part)Qwen3.5-27B features of the Table-2 folds; GPT-2-large gen-PPL cache
outputs/experiments/position_robustness/feats/ (features part)Qwen3.5-9B features of the failure-position corpus

Feature files are float32 .npy matrices, one row per line of the text file they are named after.

License

Each text keeps the license of its source, and an edited or spliced passage follows the license of the passage it was made from. Generated text is listed with the terms of the model that produced it; none of these models restricts how its outputs may be used. Everything else in this dataset (its organization, the manifests and labels, bt_scores.csv, and the cached features) is released under CC BY 4.0.

SourceRoleLicense or terms
OpenWebText (Skylion007/openwebtext)human reference and held-out pools, counterfactual parents, prefixes and human continuationsCC0 1.0
Wikipedia via WikiText-103 (Salesforce/wikitext)counterfactual parentsCC BY-SA 3.0 and GFDL
Reddit TL;DR (trl-lib/tldr)counterfactual parentsnone stated on trl-lib/tldr; it is OpenAI's filtered subset of Webis-TLDR-17, released under CC BY 4.0
SQuAD v1.1 (rajpurkar/squad)QA-faithfulness textsCC BY-SA 4.0
MAUVE human evaluation (Pillutla et al., 2021; krishnap25/mauve-experiments)human-agreement texts and the judgments behind bt_scores.csvnone stated in the source repository
GPT-2 WebText test split (openai/gpt-2-output-dataset)human-agreement referenceMIT
GPT-2 small, medium, large, XLgenerations (Table 2, Figure 5, human agreement)MIT
MDLM (kuleshov-group/mdlm-owt), LangFlow (Continuous-Rivals-Discrete/langflow-owt)generationsApache-2.0
SEDD (louaaron/sedd-small)generationscode MIT; the checkpoint states no license
ELF (embedded-language-flows)generationsthe checkpoints state no license
Qwen3-30B-A3B, Mistral-Small-24B-Instruct-2501edited variantsApache-2.0
Qwen3.5-27B, Qwen3.5-9B; GPT-2-largecached features; gen-PPL cacheApache-2.0; MIT

Citation

bibtex
@misc{liu2026coherenceawaredistributionalevaluationopenended,
      title={Coherence-Aware Distributional Evaluation of Open-Ended Text Generation},
      author={Jinnuo Liu and Junhao Zhu and Weifeng Jiang and Haoming Liu and Hongyi Wen},
      year={2026},
      eprint={2609.34240},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2609.34240},
}