nkp2mr/f26-dataprivacy-lab2
Lab 2: Synthetic forum data 7,823 comments from 300 fictional users across 103 forum threads, adapted from SynthPAI. The original authors generated the conversations with GPT-4 agents; these are not scraped accounts or real people's histories. Use the comments to investigate privacy exposure and test protective rewrites. For example, does reading several comments together reveal more than reading one? Does removing names protect the same information as changing contextual clues?… See the full description on the dataset page: https://huggingface.co/datasets/nkp2mr/f26-dataprivacy-lab2.
Lab 2: Synthetic forum data
7,823 comments from 300 fictional users across 103 forum threads, adapted from SynthPAI. The original authors generated the conversations with GPT-4 agents; these are not scraped accounts or real people's histories.
Use the comments to investigate privacy exposure and test protective rewrites. For example, does reading several comments together reveal more than reading one? Does removing names protect the same information as changing contextual clues? Which model conclusions are supported by the text, and which are guesses? Keep experiments within this fictional dataset; do not look up or profile real people.
Files
posts.jsonl:post_id,user_id,thread_id,parent_id, andtext. Give your model these comments, not the reference profiles.profiles.jsonl: the fictional attributes used to generate each user's posts. Use these separately to check results. A profile value need not be inferable from the comments, and generated comments can contradict the profile.provenance.json: the pinned source revision, transformations, and file hashes.
Reference attributes are age, sex, city_country, birth_city_country, education, occupation, income, income_level, and relationship_status. Field names and values follow the source; sex is its synthetic persona field, not a label of a real person's identity. Income amounts retain their original free-text currencies and are not directly comparable across countries.
Load in Colab
No Hugging Face account is needed. Install huggingface_hub if necessary.
import json
from pathlib import Path
from pprint import pprint
from huggingface_hub import hf_hub_download
path = hf_hub_download(
"nkp2mr/f26-dataprivacy-lab2",
"posts.jsonl",
repo_type="dataset",
token=False,
)
posts = [json.loads(line) for line in Path(path).read_text().splitlines()]
pprint(posts[:3])Download profiles.jsonl the same way when checking your results. The Hub's train split is a file-viewing convention, not a prescribed training split. These data and labels are public, not hidden grading cases.
Source and limitations
Hanna Yukhymenko, Robin Staab, Mark Vero, and Martin Vechev. A Synthetic Dataset for Personal Attribute Inference, NeurIPS 2024. Original dataset curated and shared by SRILab, ETH Zurich. Source revision: b572595f543a51db789caddbb81a9fc4edc6c32f.
This teaching adaptation preserves comment text and reply relationships. It replaces metadata identifiers, separates persona attributes from comments, and omits usernames, generation instructions, model guesses, and human reviews from the metadata. Mentions inside comment text are unchanged. File order and IDs do not represent time. The source model's biases and generation artifacts remain; performance here is not evidence of accuracy on real people or a representative population. This package does not reproduce the paper's human-annotation evaluation.
Data are distributed under the original CC BY-NC-SA 4.0 license: retain attribution, identify changes, use noncommercially, and share adaptations under the same license. The upstream code's MIT license does not replace the dataset license. This adaptation is not endorsed by the original authors.
