wwj95/privacy-aware-memory-benchmark
Privacy-Aware Memory Benchmark Paper | Code This repository provides synthetic, multi-turn, privacy-aware conversation histories used in our SP-Mem paper. The conversations contain both private information and non-private preferences across education, finance, medical, and mental domains. Data The conversation histories are stored as JSON files, with one file per synthetic user, organized by domain. Domain Users Dialogue sessions Education 250 5,250… See the full description on the dataset page: https://huggingface.co/datasets/wwj95/privacy-aware-memory-benchmark.
Privacy-Aware Memory Benchmark
This repository provides synthetic, multi-turn, privacy-aware conversation histories used in our SP-Mem paper. The conversations contain both private information and non-private preferences across education, finance, medical, and mental domains.
Data
The conversation histories are stored as JSON files, with one file per synthetic user, organized by domain.
Evaluation queries and code are available in our GitHub repository.
Usage
Load the conversation histories for a domain (education, finance, medical, or mental):
from datasets import load_dataset
histories = load_dataset(
"wwj95/privacy-aware-memory-benchmark",
"education",
split="histories",
)
Source and License
Medical profile fields were seeded from the Diseases_Symptoms dataset.
This dataset is licensed under CC BY 4.0.
Citation
@misc{wang2026whatrememberreveal,
title={What to Remember, What to Reveal: Privacy-Aware Memory for Conversational Agents},
author={Wenjie Wang and Wenhe Si and Xinyue Xu and Yue Xu},
year={2026},
eprint={2608.16551},
archivePrefix={arXiv}
}