datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bayesian-llm-safety-inference
Bayesian Latent Safety-Trait Dataset
Summary
This dataset supports Bayesian latent-trait analysis of language-model safety behavior.
It contains 90 benchmark-derived roots, three matched prompt variants per root, responses
from four target models over five runs, two independent LLM ratings per response, and one
human rating for a stratified 540-response calibration subset.
The three dimensions are harmful compliance, sycophancy, and agentic protocol violation.… See the full description on the dataset page: https://huggingface.co/datasets/Charly-X/bayesian-llm-safety-inference.p2-etf-bayesian-nonparametric-hdp-resultsrepro-on-regret-bounds-of-thompson-sampling-for-bayesian-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
bayesian-social-deduction
Bayesian Social Deduction Dataset
Project Page | Arxiv | Github
Dataset Description
This dataset contains a collection of game logs from Avalon social deduction games, generated for the "Bayesian Social Deduction with Graph-Informed Language Models" paper. The dataset includes games played by various agents, including humans, and different AI models, providing a rich resource for analyzing strategic communication, deception, and cooperation.
The dataset is organized into… See the full description on the dataset page: https://huggingface.co/datasets/shahabrahimirad/bayesian-social-deduction.
