Team Ai
Datasetpublic

wwj95/privacy-aware-memory-benchmark

Privacy-Aware Memory Benchmark Paper | Code This repository provides synthetic, multi-turn, privacy-aware conversation histories used in our SP-Mem paper. The conversations contain both private information and non-private preferences across education, finance, medical, and mental domains. Data The conversation histories are stored as JSON files, with one file per synthetic user, organized by domain. Domain Users Dialogue sessions Education 250 5,250… See the full description on the dataset page: https://huggingface.co/datasets/wwj95/privacy-aware-memory-benchmark.

sourceHugging Facecc-by-4.0updated 3d agoView on Hugging Face
0likes116downloads
Dataset Card

Privacy-Aware Memory Benchmark

Paper | Code

This repository provides synthetic, multi-turn, privacy-aware conversation histories used in our SP-Mem paper. The conversations contain both private information and non-private preferences across education, finance, medical, and mental domains.

Data

The conversation histories are stored as JSON files, with one file per synthetic user, organized by domain.

DomainUsersDialogue sessions
Education2505,250
Finance2505,250
Medical2505,250
Mental2505,250
Total1,00021,000

Evaluation queries and code are available in our GitHub repository.

Usage

Load the conversation histories for a domain (education, finance, medical, or mental):

python
from datasets import load_dataset

histories = load_dataset(
    "wwj95/privacy-aware-memory-benchmark",
    "education",
    split="histories",
)

Source and License

Medical profile fields were seeded from the Diseases_Symptoms dataset.

This dataset is licensed under CC BY 4.0.

Citation

bibtex
@misc{wang2026whatrememberreveal,
  title={What to Remember, What to Reveal: Privacy-Aware Memory for Conversational Agents},
  author={Wenjie Wang and Wenhe Si and Xinyue Xu and Yue Xu},
  year={2026},
  eprint={2608.16551},
  archivePrefix={arXiv}
}