Team Ai
Datasetpublic

jamesding0302/memgen-annotations

MemGen Annotations This is the annotation dataset for the paper How Well Does Generative Recommendation Generalize?. The annotations categorize evaluation instances under the leave-one-out protocol: test split uses the last item in the user history sequence as target, val split uses the second-to-last item as target. Columns sample_id: row index within the split in the original dataset. user_id: raw user identifier (join key). master: one of memorization… See the full description on the dataset page: https://huggingface.co/datasets/jamesding0302/memgen-annotations.

sourceHugging Faceupdated 7mo agoView on Hugging Face
1likes76downloads
Dataset Card

MemGen Annotations

This is the annotation dataset for the paper [How Well Does Generative Recommendation Generalize?](https://huggingface.co/papers/2603.19809).

<a href="https://huggingface.co/papers/2603.19809"><img src="https://img.shields.io/badge/Paper-ArXiv-red"></a> <a href="https://github.com/Jamesding000/MemGen-GR"><img src="https://img.shields.io/badge/Code-GitHub-green"></a> <a href="https://huggingface.co/jamesding0302/memgen-checkpoints"><img src="https://img.shields.io/badge/Models-Hugging%20Face-blue"></a>

The annotations categorize evaluation instances under the leave-one-out protocol:

  • —test split uses the last item in the user history sequence as target,
  • —val split uses the second-to-last item as target.

Columns

  • —sample_id: row index within the split in the original dataset.
  • —user_id: raw user identifier (join key).
  • —master: one of memorization, generalization, uncategorized.
  • —subcategories: list of {rule, hop} for fine-grained generalization types.
  • —all_labels: all string labels (e.g., ["generalization", "symmetry_3"]).

Load in M&G annotations

python
from datasets import load_dataset

labels = load_dataset(
    "jamesding0302/memgen-annotations",
    "AmazonReviews2014-Beauty",
    split="test",
)
print(labels[0])

Merge with processed dataset

python
# 1) Load your processed dataset split (must be aligned with labels by row order)
ds = pipeline.split_datasets["test"]

# 2) Append label columns to the original dataset
ds = (ds
      .add_column("master", labels["master"])
      .add_column("subcategories", labels["subcategories"])
      .add_column("all_labels", labels["all_labels"]))