Team Ai
Datasetpublic

Kkryptonite/CUE-Mem

🧠 CUE-Mem Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations 多模态对话隐式线索驱动的长期用户记忆评测基准 Text · Image · Audio | Explicit & Implicit Cues | Long-Term User Memory English · 中文说明 · Quick Start / 快速开始 · Citation / 引用 From visible routines to subtle clues: a cat bowl, a cat tree, and a background meow jointly suggest that the user has a cat. 从日常活动到隐式线索:猫碗、猫爬架与背景中的猫叫声,共同指向用户养猫这一信息。 English Overview What… See the full description on the dataset page: https://huggingface.co/datasets/Kkryptonite/CUE-Mem.

sourceHugging Facemitupdated 6d agoView on Hugging Face
0likes1.8kdownloads
Dataset Card

<div align="center">

🧠 CUE-Mem

Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations

多模态对话隐式线索驱动的长期用户记忆评测基准

![Paper](https://arxiv.org/abs/2609.32574) ![Code](https://github.com/yulinlp/CUE-MEM) ![Demo](https://reichenbach1854-hash.github.io/CUE-Mem/)

Text · Image · Audio | Explicit & Implicit Cues | Long-Term User Memory

English · 中文说明 · Quick Start / 快速开始 · Citation / 引用

</div>

[image]

From visible routines to subtle clues: a cat bowl, a cat tree, and a background meow jointly suggest that the user has a cat. 从日常活动到隐式线索:猫碗、猫爬架与背景中的猫叫声,共同指向用户养猫这一信息。


English

Overview

What should an agent remember when a user never says it directly? A recurring object in a photograph or a background sound in a voice message may carry information that matters in a later conversation.

CUE-Mem evaluates long-term user memory across text, images, and audio, covering both explicit statements and implicit cues. See the paper for benchmark design and evaluation protocols.

This repository distributes conversation histories, questions, media, and metadata. Construction and evaluation code lives on GitHub; the interactive demo presents selected examples.

Dataset at a glance

Counts below are computed from this release's metadata. A conversation record means one row in dialogue_turns.jsonl, not an individual speaker message. Sessions are counted using profile and session ID together.

ComponentCount
User profiles20
Conversation sessions648
Conversation records5,798
Evaluation questions2,674
Images4,876
Audio files2,228
Total media files7,104

Dialogue and question text is primarily Chinese; some descriptive fields are in English. Bilingual documentation does not imply parallel translations of every example.

<p align="center"> <a href="figure/datasetstatistics.pdf"><img src="figure/datasetstatistics.png" alt="Question distribution by task, modality, and evidence type / 按任务、模态与证据类型划分的问题分布" width="640"></a> </p>

Question distribution in the paper. T / I / A denote text, image, and audio; Exp. / Imp. denote explicit and implicit evidence. 论文中的问题分布:T / I / A 分别表示文本、图像和音频,Exp. / Imp. 分别表示显式和隐式证据。[View PDF / 查看原始 PDF](figure/dataset_statistics.pdf).

Tasks

The release's point field distinguishes question categories. Image-based questions can supply visual answer options through option_images.

TaskFocus`point` valuesQuestions
Entity RecallUser-related entities and attributesentity_text, entity_img1,128
Long PatternPersistent preferences and habitspref_text, pref_img363
Personalized RecommendationApply user memory to recommendationsrec_text, rec_img364
Answer RefusalInsufficient evidence or misleading premisesadversarial_text819

![Examples of the four memory tasks / 四类记忆任务示例](figure/task_example.pdf)

Four tasks, illustrated with questions, answer choices, and supporting memory clues. Open the PDF to inspect the details. 四类任务的具体示例,包括问题、候选答案与记忆线索。[View PDF / 查看原始 PDF](figure/task_example.pdf).

Label detail: qa_type is heterogeneous: some values are explicit / implicit, others are entity categories (Relationship, Pets, Items), and a few are empty. Do not use it as a universal binary evidence label. Preserve optional fields such as entity_explicitness and follow the evaluation code's interpretation.

Construction Pipeline / 数据构建流程

[image]

Construction pipeline: profile and event planning → multimodal dialogue synthesis → automatic checks and human review. 构建流程:用户档案与事件规划 → 多模态对话合成 → 自动检查与人工审核。

Files and annotations

Files are published directly at the repository root, without an extra CUE-Mem-Benchmark/ wrapper or archive extraction step. See the shared directory tree below.

Each history_with_qa_p*.json contains character_profile, multi_session_dialogues, and human-annotated QAs. The flattened JSONL files are convenient entry points for analysis.

Metadata fileContentsUseful fields
dialogue_turns.jsonlOne conversation record per linep_id, session_id, round, date, task_id, user, assistant, input_image, input_voice_message
questions.jsonlOne question per lineqa_id, p_id, question, answer, point, qa_type, clue, option_images, source_file
media_manifest.jsonlMedia inventorypath, media_type, format, size_bytes
checksums.sha256SHA-256 checksums for data/ filesHash and relative file path

Fields vary by modality and task. Audio-only turns may have no user text field. Captions are available in image_caption, voice_caption, and user_voice_message_caption where present.

Media paths, source_file, and manifest path use / separators and are relative to the dataset root. However, clue can contain evidence identifiers such as D09:02 or D09-001.wav, which are not necessarily repository paths. Include p_id when joining session or turn references across profiles.

Evaluation notes

Keep reference answers, rationales, and evidence annotations outside the evaluated system's memory unless the protocol explicitly provides them, such as an oracle-evidence baseline. State whether the system receives raw media, captions, or both. Preserve profile/session boundaries and conversation order; report the dataset revision, question selection, model, and memory configuration.

Generated scenarios and profile media should not be interpreted as verified observations about real people. Consult the paper for construction details and limitations, and the GitHub README for environment setup, dataset preparation, and RQ1–RQ3 evaluation instructions.


中文说明

基准简介

用户没有直接说出口的信息,智能体能否在之后的对话中记住并使用?照片中反复出现的物品、语音里的环境声音,都可能是理解用户的重要线索。

CUE-Mem 面向文本、图像和音频交织的长期对话,评测记忆系统对显式信息与隐式线索的保留、检索和利用能力。基准设计与评测协议见论文。

本仓库提供对话历史、问答、媒体文件及元数据;构建与评测代码位于 GitHub,可通过在线 Demo浏览精选示例。

数据规模与任务

当前发布数据包含 20 个用户档案、648 个会话、5,798 条对话记录、2,674 道问题,以及 4,876 张图像、2,228 个音频文件。会话按用户和会话 ID 联合计数;对话记录数对应 dialogue_turns.jsonl 的行数,并非将双方消息分别计数。

任务评测目标`point` 字段题数
实体回忆(Entity Recall)记住与用户相关的实体及属性entity_text、entity_img1,128
长期模式(Long Pattern)从多次交互中归纳偏好和习惯pref_text、pref_img363
个性化推荐(Personalized Recommendation)将用户记忆用于新的推荐问题rec_text、rec_img364
拒答(Answer Refusal)识别证据不足或问题前提不可靠的情况adversarial_text819

对话和问题主要为中文,部分描述字段为英文;双语说明并不代表全部样本都有平行翻译。图像类问题可通过 option_images 提供图片选项。

文件与字段

仓库根目录直接包含 README.md、LICENSE、data/ 和 metadata/,没有额外套一层目录,下载后无需拼接或解压分卷。完整目录树见下方。

data/dialog/base/ 中每个用户对应一个 JSON,包含 character_profile(用户档案)、multi_session_dialogues(多会话对话)和 human-annotated QAs(问答)。data/event/ 保存对话图像与音频,data/qa/ 保存问题图片,data/profile/generated_portraits/ 保存生成的用户/实体参考图像。

元数据文件内容与读取要点
dialogue_turns.jsonl每行一条对话记录;用 p_id、session_id、round 定位,媒体见 input_image、input_voice_message
questions.jsonl每行一道问题;包含问题、答案、任务类型、线索及部分题目的图片选项
media_manifest.jsonl媒体清单:path、media_type、format、size_bytes
checksums.sha256data/ 文件的 SHA-256 校验值,用于检查下载完整性

不同模态和题型的字段不完全一致。音频轮次可能没有 user 文本字段;描述可见 image_caption、voice_caption、user_voice_message_caption 等可选字段。

读取时请注意:

  • —媒体路径、source_file 和清单中的 path 相对于数据集根目录;clue 中的 D09:02、D09-001.wav 等可能是线索 ID,不能一律当作磁盘路径。跨用户关联时应同时使用 p_id。
  • —qa_type 并非统一的显式/隐式二分类:部分为 explicit 或 implicit,实体题也使用 Relationship、Pets、Items,另有少量空值。请结合 entity_explicitness 等字段及评测代码解释标签。

评测建议

请保持用户、会话边界和原始对话顺序,并记录数据版本、抽样方式、模型及记忆系统配置。除协议明确允许的情况(如 oracle-evidence 基线)外,不应将标准答案、解释或标注线索提前写入被评测系统的记忆。使用原始媒体、文本描述或二者组合时,也应在结果中说明。

生成场景及用户媒体不应被视为真实人物的已核实经历。构建方法和局限性见论文;环境配置、数据准备及 RQ1–RQ3 实验步骤见 GitHub README。


Repository Layout / 仓库结构

text
.
├── README.md
├── LICENSE
├── data/
│   ├── dialog/base/                  # Profile histories + QA / 用户历史与问答
│   │   ├── history_with_qa_p0.json
│   │   └── ...
│   ├── event/images/                 # Event images / 对话事件图像
│   ├── event/voice_mixed/            # Event audio / 对话音频
│   ├── qa/pref_images/               # Preference / 偏好题图片
│   ├── qa/rec_images/                # Recommendation / 推荐题图片
│   ├── qa/entity_images/             # Entity / 实体题图片
│   └── profile/generated_portraits/  # Profile/entity references / 档案与实体参考图
└── metadata/
    ├── dialogue_turns.jsonl
    ├── questions.jsonl
    ├── media_manifest.jsonl
    └── checksums.sha256

<a id="quick-start"></a>

Quick Start / 快速开始

1. Download / 下载

Download the full release with the Hugging Face CLI. / 使用 Hugging Face CLI 下载完整数据:

bash
python -m pip install -U huggingface_hub
hf download Kkryptonite/CUE-Mem --repo-type dataset --local-dir benchmark-data

For a lightweight first look, download only metadata and profile JSON files. Download media before multimodal evaluation. / 若只想先查看结构,可仅下载元数据与用户 JSON;多模态评测前需补齐媒体:

python
from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="Kkryptonite/CUE-Mem",
    repo_type="dataset",
    local_dir="benchmark-data",
    allow_patterns=["README.md", "LICENSE", "metadata/*", "data/dialog/base/*.json"],
)

For reproducibility, set revision to a commit hash. See the download guide. / 固定版本时可设置 revision 为 commit hash,详见下载说明。

2. Read records / 读取数据

python
import json
from pathlib import Path

root = Path("benchmark-data")

def read_jsonl(relative_path):
    with (root / relative_path).open(encoding="utf-8") as handle:
        for line in handle:
            if line.strip():
                yield json.loads(line)

question = next(read_jsonl("metadata/questions.jsonl"))
print(question["qa_id"], question["question"])
for label, relative_path in question.get("option_images", {}).items():
    print(label, root / relative_path)

turn = next(read_jsonl("metadata/dialogue_turns.jsonl"))
print(turn["p_id"], turn["session_id"], turn["round"])
for field in ("input_image", "input_voice_message"):
    for relative_path in turn.get(field, []):
        print(field, root / relative_path)

3. Verify files / 校验完整性

Run after downloading the full data/ directory. / 完整下载 data/ 后执行:

bash
cd benchmark-data

# Linux
sha256sum -c metadata/checksums.sha256

# macOS (alternative / 替代命令)
shasum -a 256 -c metadata/checksums.sha256

License / 许可证

This repository includes an MIT License. Refer to the file for the full terms. / 本仓库附有 MIT 许可证,完整条款以该文件为准。

Citation

If you use CUE-Mem in your research, please cite the paper. / 如在研究中使用 CUE-Mem,请引用以下论文:

bibtex
@misc{hu2026cuemem,
  title={CUE-Mem: Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations},
  author={Yulin Hu and Yanyan Zhao and Zimo Long and Xing Fu and Mengtong Ji and Weixiang Zhao and Yutai Hou and Qianchao Wang and Dandan Tu},
  year={2026},
  eprint={2609.32574},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  url={https://arxiv.org/abs/2609.32574}
}

Code and evaluation questions: GitHub Issues. Dataset questions: Hugging Face Discussions. / 代码与评测问题请提交至 GitHub Issues,数据集问题可在 Hugging Face Discussions 反馈。