yaelwalker/emotion-probe-controls
Bodily-state stories Synthetic fictional stories for extracting bodily-state vectors from large language models. Each story portrays a character experiencing a labeled bodily feeling or physiological state in a given scenario. Files and schema File Rows Coverage Generation model body-state/stories.parquet 10,400 104 states × 100 topics × 1 story per combination google/gemini-3.8-flash The Parquet file has exactly three columns, all strings: state:… See the full description on the dataset page: https://huggingface.co/datasets/yaelwalker/emotion-probe-controls.
Bodily-state stories
Synthetic fictional stories for extracting bodily-state vectors from large language models. Each story portrays a character experiencing a labeled bodily feeling or physiological state in a given scenario.
Files and schema
The Parquet file has exactly three columns, all strings:
- `state`: the target bodily-state label, e.g.
hungry,in pain, orcold. - `topic`: the scenario/premise, using the original emotion-probes topics.
- `story`: the generated story, preserved verbatim from the JSONL source.
This matches the layout of ryancodrai/emotion-probes's expression/stories.parquet, except that the class column is named state instead of emotion. Labels are strings, not integer-encoded classes. Rows are sorted by state, topic, and the source story index. Generation bookkeeping is not included as columns.
Loading
Replace YOUR_USERNAME/YOUR_DATASET with the uploaded dataset repository ID:
from datasets import load_dataset
stories = load_dataset(
"YOUR_USERNAME/YOUR_DATASET",
data_files="body-state/stories.parquet",
)["train"]
labels = stories["state"] # Previously stories["emotion"]
texts = stories["story"]If this README is uploaded as the dataset repository's root README.md, the configuration above also supports:
stories = load_dataset("YOUR_USERNAME/YOUR_DATASET", name="body-state")["train"]Future story types can be added in separate directories and named configurations without combining incompatible Parquet schemas into one split. Explicitly selecting data_files remains supported.
No neutral-story dataset is included. Extraction pipelines that use the original neutral stories for confound removal must continue loading those from ryancodrai/emotion-probes (or supply a separately generated neutral set).
Generation and limitations
Stories were generated with google/gemini-3.8-flash via OpenRouter. The prompt adapts the original emotion-story task to bodily feelings, focusing on concrete physical sensations, bodily processes, behavior, and their interaction with the scenario. It requests independent one-paragraph stories, first- or third-person narration, and avoidance of the target state word and direct synonyms.
This is synthetic data. Target-state fidelity, synonym avoidance, emotional confounds, and semantic diversity have not been independently verified. Some states overlap in meaning.
Source attribution
The 100 topics and starting expression-story generation prompt come from Ryan Codrai's Emotion Probes Dataset, licensed CC BY 4.0. The generation prompt was modified to portray bodily states rather than emotions. The stories in this file were newly generated, not copied from the original dataset.
