Kkryptonite/CUE-Mem
🧠 CUE-Mem Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations 多模态对话隐式线索驱动的长期用户记忆评测基准 Text · Image · Audio | Explicit & Implicit Cues | Long-Term User Memory English · 中文说明 · Quick Start / 快速开始 · Citation / 引用 From visible routines to subtle clues: a cat bowl, a cat tree, and a background meow jointly suggest that the user has a cat. 从日常活动到隐式线索:猫碗、猫爬架与背景中的猫叫声,共同指向用户养猫这一信息。 English Overview What… See the full description on the dataset page: https://huggingface.co/datasets/Kkryptonite/CUE-Mem.
<div align="center">
🧠 CUE-Mem
Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations
多模态对话隐式线索驱动的长期用户记忆评测基准
  
Text · Image · Audio | Explicit & Implicit Cues | Long-Term User Memory
English · 中文说明 · Quick Start / 快速开始 · Citation / 引用
</div>
From visible routines to subtle clues: a cat bowl, a cat tree, and a background meow jointly suggest that the user has a cat. 从日常活动到隐式线索:猫碗、猫爬架与背景中的猫叫声,共同指向用户养猫这一信息。
English
Overview
What should an agent remember when a user never says it directly? A recurring object in a photograph or a background sound in a voice message may carry information that matters in a later conversation.
CUE-Mem evaluates long-term user memory across text, images, and audio, covering both explicit statements and implicit cues. See the paper for benchmark design and evaluation protocols.
This repository distributes conversation histories, questions, media, and metadata. Construction and evaluation code lives on GitHub; the interactive demo presents selected examples.
Dataset at a glance
Counts below are computed from this release's metadata. A conversation record means one row in dialogue_turns.jsonl, not an individual speaker message. Sessions are counted using profile and session ID together.
Dialogue and question text is primarily Chinese; some descriptive fields are in English. Bilingual documentation does not imply parallel translations of every example.
<p align="center"> <a href="figure/datasetstatistics.pdf"><img src="figure/datasetstatistics.png" alt="Question distribution by task, modality, and evidence type / 按任务、模态与证据类型划分的问题分布" width="640"></a> </p>
Question distribution in the paper. T / I / A denote text, image, and audio; Exp. / Imp. denote explicit and implicit evidence. 论文中的问题分布:T / I / A 分别表示文本、图像和音频,Exp. / Imp. 分别表示显式和隐式证据。[View PDF / 查看原始 PDF](figure/dataset_statistics.pdf).
Tasks
The release's point field distinguishes question categories. Image-based questions can supply visual answer options through option_images.

Four tasks, illustrated with questions, answer choices, and supporting memory clues. Open the PDF to inspect the details. 四类任务的具体示例,包括问题、候选答案与记忆线索。[View PDF / 查看原始 PDF](figure/task_example.pdf).
Label detail: qa_type is heterogeneous: some values are explicit / implicit, others are entity categories (Relationship, Pets, Items), and a few are empty. Do not use it as a universal binary evidence label. Preserve optional fields such as entity_explicitness and follow the evaluation code's interpretation.
Construction Pipeline / 数据构建流程
Construction pipeline: profile and event planning → multimodal dialogue synthesis → automatic checks and human review. 构建流程:用户档案与事件规划 → 多模态对话合成 → 自动检查与人工审核。
Files and annotations
Files are published directly at the repository root, without an extra CUE-Mem-Benchmark/ wrapper or archive extraction step. See the shared directory tree below.
Each history_with_qa_p*.json contains character_profile, multi_session_dialogues, and human-annotated QAs. The flattened JSONL files are convenient entry points for analysis.
Fields vary by modality and task. Audio-only turns may have no user text field. Captions are available in image_caption, voice_caption, and user_voice_message_caption where present.
Media paths, source_file, and manifest path use / separators and are relative to the dataset root. However, clue can contain evidence identifiers such as D09:02 or D09-001.wav, which are not necessarily repository paths. Include p_id when joining session or turn references across profiles.
Evaluation notes
Keep reference answers, rationales, and evidence annotations outside the evaluated system's memory unless the protocol explicitly provides them, such as an oracle-evidence baseline. State whether the system receives raw media, captions, or both. Preserve profile/session boundaries and conversation order; report the dataset revision, question selection, model, and memory configuration.
Generated scenarios and profile media should not be interpreted as verified observations about real people. Consult the paper for construction details and limitations, and the GitHub README for environment setup, dataset preparation, and RQ1–RQ3 evaluation instructions.
中文说明
基准简介
用户没有直接说出口的信息,智能体能否在之后的对话中记住并使用?照片中反复出现的物品、语音里的环境声音,都可能是理解用户的重要线索。
CUE-Mem 面向文本、图像和音频交织的长期对话,评测记忆系统对显式信息与隐式线索的保留、检索和利用能力。基准设计与评测协议见论文。
本仓库提供对话历史、问答、媒体文件及元数据;构建与评测代码位于 GitHub,可通过在线 Demo浏览精选示例。
数据规模与任务
当前发布数据包含 20 个用户档案、648 个会话、5,798 条对话记录、2,674 道问题,以及 4,876 张图像、2,228 个音频文件。会话按用户和会话 ID 联合计数;对话记录数对应 dialogue_turns.jsonl 的行数,并非将双方消息分别计数。
对话和问题主要为中文,部分描述字段为英文;双语说明并不代表全部样本都有平行翻译。图像类问题可通过 option_images 提供图片选项。
文件与字段
仓库根目录直接包含 README.md、LICENSE、data/ 和 metadata/,没有额外套一层目录,下载后无需拼接或解压分卷。完整目录树见下方。
data/dialog/base/ 中每个用户对应一个 JSON,包含 character_profile(用户档案)、multi_session_dialogues(多会话对话)和 human-annotated QAs(问答)。data/event/ 保存对话图像与音频,data/qa/ 保存问题图片,data/profile/generated_portraits/ 保存生成的用户/实体参考图像。
不同模态和题型的字段不完全一致。音频轮次可能没有 user 文本字段;描述可见 image_caption、voice_caption、user_voice_message_caption 等可选字段。
读取时请注意:
- 媒体路径、
source_file和清单中的path相对于数据集根目录;clue中的D09:02、D09-001.wav等可能是线索 ID,不能一律当作磁盘路径。跨用户关联时应同时使用p_id。 qa_type并非统一的显式/隐式二分类:部分为explicit或implicit,实体题也使用Relationship、Pets、Items,另有少量空值。请结合entity_explicitness等字段及评测代码解释标签。
评测建议
请保持用户、会话边界和原始对话顺序,并记录数据版本、抽样方式、模型及记忆系统配置。除协议明确允许的情况(如 oracle-evidence 基线)外,不应将标准答案、解释或标注线索提前写入被评测系统的记忆。使用原始媒体、文本描述或二者组合时,也应在结果中说明。
生成场景及用户媒体不应被视为真实人物的已核实经历。构建方法和局限性见论文;环境配置、数据准备及 RQ1–RQ3 实验步骤见 GitHub README。
Repository Layout / 仓库结构
.
├── README.md
├── LICENSE
├── data/
│ ├── dialog/base/ # Profile histories + QA / 用户历史与问答
│ │ ├── history_with_qa_p0.json
│ │ └── ...
│ ├── event/images/ # Event images / 对话事件图像
│ ├── event/voice_mixed/ # Event audio / 对话音频
│ ├── qa/pref_images/ # Preference / 偏好题图片
│ ├── qa/rec_images/ # Recommendation / 推荐题图片
│ ├── qa/entity_images/ # Entity / 实体题图片
│ └── profile/generated_portraits/ # Profile/entity references / 档案与实体参考图
└── metadata/
├── dialogue_turns.jsonl
├── questions.jsonl
├── media_manifest.jsonl
└── checksums.sha256<a id="quick-start"></a>
Quick Start / 快速开始
1. Download / 下载
Download the full release with the Hugging Face CLI. / 使用 Hugging Face CLI 下载完整数据:
python -m pip install -U huggingface_hub
hf download Kkryptonite/CUE-Mem --repo-type dataset --local-dir benchmark-dataFor a lightweight first look, download only metadata and profile JSON files. Download media before multimodal evaluation. / 若只想先查看结构,可仅下载元数据与用户 JSON;多模态评测前需补齐媒体:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="Kkryptonite/CUE-Mem",
repo_type="dataset",
local_dir="benchmark-data",
allow_patterns=["README.md", "LICENSE", "metadata/*", "data/dialog/base/*.json"],
)For reproducibility, set revision to a commit hash. See the download guide. / 固定版本时可设置 revision 为 commit hash,详见下载说明。
2. Read records / 读取数据
import json
from pathlib import Path
root = Path("benchmark-data")
def read_jsonl(relative_path):
with (root / relative_path).open(encoding="utf-8") as handle:
for line in handle:
if line.strip():
yield json.loads(line)
question = next(read_jsonl("metadata/questions.jsonl"))
print(question["qa_id"], question["question"])
for label, relative_path in question.get("option_images", {}).items():
print(label, root / relative_path)
turn = next(read_jsonl("metadata/dialogue_turns.jsonl"))
print(turn["p_id"], turn["session_id"], turn["round"])
for field in ("input_image", "input_voice_message"):
for relative_path in turn.get(field, []):
print(field, root / relative_path)3. Verify files / 校验完整性
Run after downloading the full data/ directory. / 完整下载 data/ 后执行:
cd benchmark-data
# Linux
sha256sum -c metadata/checksums.sha256
# macOS (alternative / 替代命令)
shasum -a 256 -c metadata/checksums.sha256License / 许可证
This repository includes an MIT License. Refer to the file for the full terms. / 本仓库附有 MIT 许可证,完整条款以该文件为准。
Citation
If you use CUE-Mem in your research, please cite the paper. / 如在研究中使用 CUE-Mem,请引用以下论文:
@misc{hu2026cuemem,
title={CUE-Mem: Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations},
author={Yulin Hu and Yanyan Zhao and Zimo Long and Xing Fu and Mengtong Ji and Weixiang Zhao and Yutai Hou and Qianchao Wang and Dandan Tu},
year={2026},
eprint={2609.32574},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2609.32574}
}Code and evaluation questions: GitHub Issues. Dataset questions: Hugging Face Discussions. / 代码与评测问题请提交至 GitHub Issues,数据集问题可在 Hugging Face Discussions 反馈。
