Team Ai
Datasetpublic

suchirsalhan/bayesian-coherence-lm

Bayesian Coherence of LMs — Prompt Sets Prompt sets for measuring the Bayesian-coherence of language models via the incoherence certificate |log R| (local) and the cross-lingual/transitive cycle ratio |log cycle| (global). A model's prompts induce an implicit joint distribution over entities and relations; a model is coherent on a quartet iff its loop of conditional inferences multiplies back to 1 (log R = 0). Code: https://github.com/suchirsalhan/bayesian-coherence-lm… See the full description on the dataset page: https://huggingface.co/datasets/suchirsalhan/bayesian-coherence-lm.

sourceHugging Facemitupdated 2d agoView on Hugging Face
0likes34downloads
Dataset Card

Bayesian Coherence of LMs — Prompt Sets

Prompt sets for measuring the Bayesian-coherence of language models via the incoherence certificate |log R| (local) and the cross-lingual/transitive cycle ratio |log cycle| (global). A model's prompts induce an implicit joint distribution over entities and relations; a model is coherent on a quartet iff its loop of conditional inferences multiplies back to 1 (log R = 0).

Code: https://github.com/suchirsalhan/bayesian-coherence-lm

The certificate

For a two-slot template with base values x_i, x_j and alternatives x'_i, x'_j:

        p(x_i|x_j) p(x_j|x'_i) p(x'_i|x'_j) p(x'_j|x_i)
R  =  ─────────────────────────────────────────────────
        p(x_j|x_i) p(x_i|x'_j) p(x'_j|x'_i) p(x'_i|x_j)

Every conditional is a single-token forced choice read from logits, so the eval is cheap. Two backends share the certificate: an MLM path (fill one slot, [MASK] the other) and a causal path (forward/reverse templates, continuation log-prob) — the latter handles multi-token entities and is primary.

Configs

configrowswhat it is
pilot_prompts160The exact hand-written prompt set used by the pilot experiments, materialised into concrete prompt strings (Exp1 factual, Exp1b synthetic micro-worlds, Exp2 generics, Exp3 multilingual, and the prompt-stability paraphrase family). Both mlm and causal backends.
microworld100,000Exp1b at scale — nonce entities with a known latent joint (contamination-free).
multilingual100,000Exp3 at scale — capital-of quartets realised per language.
multilingual_cycles2,000Cross-lingual cycles (en→fr→de→es→it→pt→nl) for the global `log cycle`.
flores3,000Translation-equivalent quartets from FLORES parallel text.
flores_cycles50012-language translation cycles.
pararel8,000ParaRel-style paraphrase templates per relation (prompt-stability at scale).
values100,000Exp4 value consistency — country×trait preference quartets (subjective vs objective traits).

The pilot_prompts config is self-contained and corresponds exactly to the runs reported in the repo's PILOT_NOTES.md. The other configs are the 100K-scale datasets consumed by pilots/run_dataset.py for the model sweep.

Schemas

`pilot_prompts` — one row per materialised prompt:

json
{"id": "pilot_0000", "experiment": "exp1", "backend": "mlm",
 "quartet": "capital_fr_de", "lang": "en", "predict_slot": "I",
 "prompt": "[MASK] is the capital of France .",
 "candidate_base": "Paris", "candidate_alt": "Berlin",
 "template": "{I} is the capital of {J} .",
 "xi": "Paris", "xip": "Berlin", "xj": "France", "xjp": "Germany"}

For backend: causal, rows carry fwd_tmpl/rev_tmpl and a direction (forward/reverse) instead of template; prompt is the filled stem and the model scores candidate_base vs candidate_alt as the continuation.

Quartet configs (microworld, multilingual, pararel, flores) — one row per quartet:

json
{"id": "ml_en_0", "dataset": "multilingual", "relation": "capital_of",
 "lang": "en", "fwd_tmpl": "The capital of {J} is",
 "rev_tmpl": "The city {I} is the capital of",
 "xi": " Kabul", "xip": " Tirana", "xj": " Afghanistan", "xjp": " Albania"}

Cycle configs (multilingual_cycles, flores_cycles) — one row per cycle, carrying langs and an edges list of per-language quartets.

`values` — lead, two countries c1/c2, two traits, and the four materialised comparison statements stmt_c1_t1 … stmt_c2_t2.

Usage

python
from datasets import load_dataset

pilots = load_dataset("suchirsalhan/bayesian-coherence-lm", "pilot_prompts")
ml     = load_dataset("suchirsalhan/bayesian-coherence-lm", "multilingual", split="train")

Citation

Operationalises the coherence theory of Emerson (2025). If you use these prompt sets, please cite that work and this repository.