datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodel-capitulation-interp
Multi-model wrongful-capitulation internal-readout dataset
Per-turn internal readouts + behavioral labels from two-model collaborative conversations (Qwen2.5-3B-Instruct × gemma-2-2b-it) on 6 reasoning benchmarks, restricted to the disagreement subset (one model right, one wrong solo). Built to test whether a linear correctness probe on the residual stream can predict wrongful capitulation (a model abandoning an answer it knew was correct under a partner's wrong assertion)… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/multimodel-capitulation-interp.multi-model-cot-missing-answers
Multi-model CoT responses without a final answer
Teacher responses from the text part of the next_jev stage 1 data (built from JonesLin/multi-model-cot-2730)
whose final_answer is empty, collected so they can be re-run with a model.
One row per response: 23,615 train and 500 validation rows.
The answer_and_cot stage 1 objective trains the backbone to emit the final answer, so it needs an answer
for every response. These rows have none.
Why the answer is missing… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/multi-model-cot-missing-answers.
