suchirsalhan/bayesian-coherence-lm
Bayesian Coherence of LMs — Prompt Sets Prompt sets for measuring the Bayesian-coherence of language models via the incoherence certificate |log R| (local) and the cross-lingual/transitive cycle ratio |log cycle| (global). A model's prompts induce an implicit joint distribution over entities and relations; a model is coherent on a quartet iff its loop of conditional inferences multiplies back to 1 (log R = 0). Code: https://github.com/suchirsalhan/bayesian-coherence-lm… See the full description on the dataset page: https://huggingface.co/datasets/suchirsalhan/bayesian-coherence-lm.
Bayesian Coherence of LMs — Prompt Sets
Prompt sets for measuring the Bayesian-coherence of language models via the incoherence certificate |log R| (local) and the cross-lingual/transitive cycle ratio |log cycle| (global). A model's prompts induce an implicit joint distribution over entities and relations; a model is coherent on a quartet iff its loop of conditional inferences multiplies back to 1 (log R = 0).
Code: https://github.com/suchirsalhan/bayesian-coherence-lm
The certificate
For a two-slot template with base values x_i, x_j and alternatives x'_i, x'_j:
p(x_i|x_j) p(x_j|x'_i) p(x'_i|x'_j) p(x'_j|x_i)
R = ─────────────────────────────────────────────────
p(x_j|x_i) p(x_i|x'_j) p(x'_j|x'_i) p(x'_i|x_j)Every conditional is a single-token forced choice read from logits, so the eval is cheap. Two backends share the certificate: an MLM path (fill one slot, [MASK] the other) and a causal path (forward/reverse templates, continuation log-prob) — the latter handles multi-token entities and is primary.
Configs
The pilot_prompts config is self-contained and corresponds exactly to the runs reported in the repo's PILOT_NOTES.md. The other configs are the 100K-scale datasets consumed by pilots/run_dataset.py for the model sweep.
Schemas
`pilot_prompts` — one row per materialised prompt:
{"id": "pilot_0000", "experiment": "exp1", "backend": "mlm",
"quartet": "capital_fr_de", "lang": "en", "predict_slot": "I",
"prompt": "[MASK] is the capital of France .",
"candidate_base": "Paris", "candidate_alt": "Berlin",
"template": "{I} is the capital of {J} .",
"xi": "Paris", "xip": "Berlin", "xj": "France", "xjp": "Germany"}For backend: causal, rows carry fwd_tmpl/rev_tmpl and a direction (forward/reverse) instead of template; prompt is the filled stem and the model scores candidate_base vs candidate_alt as the continuation.
Quartet configs (microworld, multilingual, pararel, flores) — one row per quartet:
{"id": "ml_en_0", "dataset": "multilingual", "relation": "capital_of",
"lang": "en", "fwd_tmpl": "The capital of {J} is",
"rev_tmpl": "The city {I} is the capital of",
"xi": " Kabul", "xip": " Tirana", "xj": " Afghanistan", "xjp": " Albania"}Cycle configs (multilingual_cycles, flores_cycles) — one row per cycle, carrying langs and an edges list of per-language quartets.
`values` — lead, two countries c1/c2, two traits, and the four materialised comparison statements stmt_c1_t1 … stmt_c2_t2.
Usage
from datasets import load_dataset
pilots = load_dataset("suchirsalhan/bayesian-coherence-lm", "pilot_prompts")
ml = load_dataset("suchirsalhan/bayesian-coherence-lm", "multilingual", split="train")Citation
Operationalises the coherence theory of Emerson (2025). If you use these prompt sets, please cite that work and this repository.
