Team Ai
Datasetpublic

yaelwalker/emotion-probe-controls

Bodily-state stories Synthetic fictional stories for extracting bodily-state vectors from large language models. Each story portrays a character experiencing a labeled bodily feeling or physiological state in a given scenario. Files and schema File Rows Coverage Generation model body-state/stories.parquet 10,400 104 states × 100 topics × 1 story per combination google/gemini-3.8-flash The Parquet file has exactly three columns, all strings: state:… See the full description on the dataset page: https://huggingface.co/datasets/yaelwalker/emotion-probe-controls.

sourceHugging Faceupdated 17h agoView on Hugging Face
0likes
Dataset Card

Bodily-state stories

Synthetic fictional stories for extracting bodily-state vectors from large language models. Each story portrays a character experiencing a labeled bodily feeling or physiological state in a given scenario.

Files and schema

FileRowsCoverageGeneration model
body-state/stories.parquet10,400104 states × 100 topics × 1 story per combinationgoogle/gemini-3.8-flash

The Parquet file has exactly three columns, all strings:

  • —`state`: the target bodily-state label, e.g. hungry, in pain, or cold.
  • —`topic`: the scenario/premise, using the original emotion-probes topics.
  • —`story`: the generated story, preserved verbatim from the JSONL source.

This matches the layout of ryancodrai/emotion-probes's expression/stories.parquet, except that the class column is named state instead of emotion. Labels are strings, not integer-encoded classes. Rows are sorted by state, topic, and the source story index. Generation bookkeeping is not included as columns.

Loading

Replace YOUR_USERNAME/YOUR_DATASET with the uploaded dataset repository ID:

python
from datasets import load_dataset

stories = load_dataset(
    "YOUR_USERNAME/YOUR_DATASET",
    data_files="body-state/stories.parquet",
)["train"]

labels = stories["state"]  # Previously stories["emotion"]
texts = stories["story"]

If this README is uploaded as the dataset repository's root README.md, the configuration above also supports:

python
stories = load_dataset("YOUR_USERNAME/YOUR_DATASET", name="body-state")["train"]

Future story types can be added in separate directories and named configurations without combining incompatible Parquet schemas into one split. Explicitly selecting data_files remains supported.

No neutral-story dataset is included. Extraction pipelines that use the original neutral stories for confound removal must continue loading those from ryancodrai/emotion-probes (or supply a separately generated neutral set).

Generation and limitations

Stories were generated with google/gemini-3.8-flash via OpenRouter. The prompt adapts the original emotion-story task to bodily feelings, focusing on concrete physical sensations, bodily processes, behavior, and their interaction with the scenario. It requests independent one-paragraph stories, first- or third-person narration, and avoidance of the target state word and direct synonyms.

This is synthetic data. Target-state fidelity, synonym avoidance, emotional confounds, and semantic diversity have not been independently verified. Some states overlap in meaning.

Source attribution

The 100 topics and starting expression-story generation prompt come from Ryan Codrai's Emotion Probes Dataset, licensed CC BY 4.0. The generation prompt was modified to portray bodily states rather than emotions. The stories in this file were newly generated, not copied from the original dataset.