datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
longmemeval-cleanedThis dataset replaces the original LongMemEval dataset. The main difference is that this version removes noisy history sessions that interfere with the answer correctness. More detailed session processing information can be found here.
longmemeval⚠️ This dataset is deprecated. It is replaced by longmemeval-cleaned (https://huggingface.co/datasets/xiaowu0162/longmemeval-cleaned) which noisy history sessions that interfere with the answer correctness.
longmemeval-v2
LongMemEval-V2 Data
Project Page | Paper | GitHub
LongMemEval-V2 (LME-V2) is an evaluation benchmark for long-term memory in web and enterprise agents. It contains 451 manually curated questions and 1,870 task trajectories drawn from customized WebArena-style and ServiceNow-style environments.
The benchmark evaluates five core memory abilities:
Static state recall: remembers important landmarks and page layouts.
Dynamic state tracking: understands how states change over time.… See the full description on the dataset page: https://huggingface.co/datasets/xiaowu0162/longmemeval-v2.funes-xiaowu0162-longmemeval-cleaned-s
Funes recall store — LongMemEval_s cleaned corpus
A funes recall store built by indexing the
longmemeval_s_cleaned.json haystack of
xiaowu0162/longmemeval-cleaned
(LongMemEval, arXiv:2410.10813) — every unique
chat session across all 500 questions' haystacks, in one corpus-wide store.
What this is
This is not a raw trace dataset — it is a pre-built funes index: the source
sessions chunked into content blocks and embedded, stored as a
Lance table (chunks.lance).… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-xiaowu0162-longmemeval-cleaned-s.longmemeval-s-cleanedLongMemEvalDocLongMemEval_24k
Datasets for Sliding Window Attention Adaptation (SWAA)
This repository contains evaluation datasets for Large Language Models (LLMs) adapted with Sliding Window Attention Adaptation (SWAA), as presented in the paper Sliding Window Attention Adaptation.
The associated code for SWAA implementation and evaluation scripts can be found on GitHub: https://github.com/yuyijiong/sliding-window-attention-adaptation
This repository specifically hosts evaluation datasets:
longmemeval_24k: A… See the full description on the dataset page: https://huggingface.co/datasets/yuyijiong/LongMemEval_24k.LongMemEval-LoCoMo-distilledlongmemeval-zh
LongMemEval-S 中文版(LongMemEval-S Chinese)
LongMemEval-S 的全量中文化版本,用于评测长对话记忆系统。源数据集采用 MIT 许可,本译本同样以 MIT 发布。
翻译模型:智谱 glm-5.3-flash,全局术语表统一人名/地名译法(新华社音译)
结构与源数据完全一致(消息数、role 顺序、session id、日期逐条对齐)
100% 全量校验:结构 0 错、无漏译(226,773 条消息逐条扫描)、人名拉丁残留 0
文件
文件
说明
longmemeval_s_cleaned_zh.json
中文版主数据集(194MB,470 题)
longmemeval_s_cleaned.json
英文原始数据集(265MB,500 题,未改动)
pruned_questions.json
剔除清单(30 题:question_id / 题目 / 类型)
与源数据的差异(对账:470 保留 + 30 剔除 =… See the full description on the dataset page: https://huggingface.co/datasets/justis-xu/longmemeval-zh.LongMemEval
LongMemEval
An MTEB dataset
Massive Text Embedding Benchmark
LMEB dialogue-memory retrieval task based on LongMemEval, evaluating single-session, multi-session, preference, knowledge-update, and temporal queries.
Task category
Retrieval (text-to-text)
Domains
Social, Spoken
Reference
LMEB: Long-horizon Memory Embedding Benchmark
Source datasets:
KaLM-Embedding/LMEB
How to evaluate on this task
You can evaluate an embedding model on this dataset using… See the full description on the dataset page: https://huggingface.co/datasets/mteb/LongMemEval.longmemeval-com-dom-per-question
longmemeval-com-dom-per-question
Private archive of per_question.jsonl rows from the LME-v2 (LongMemEval COM/DOM) eval runs
on Viraj's Windows box (C:/Users/12066/Desktop/longmemeval-com-dom-memory/runs/).
These are the last single-copy artifacts of that eval arc -- gitignored under runs/ per the
repo's heavy-data-goes-to-HF rule, and this box is OOM-prone, so they live here as the backup copy.
Status
BLOCKED as of 2026-07-04: the Vi0509 HF account's private… See the full description on the dataset page: https://huggingface.co/datasets/Vi0509/longmemeval-com-dom-per-question.cleaned-longmemeval-s
Cleaned Version of LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
This dataset is a cleaned version of LongMemEval by Wu et al. (2024).
This dataset was used in Context Rot.
Modifications
Removed ambiguous question-answers
Fixed focused (orcale) version to ensure question can be fully answered by input
Citation
If you use this dataset, please cite the original authors:
@article{wu2024longmemeval,
title={LongMemEval:… See the full description on the dataset page: https://huggingface.co/datasets/kellyhongg/cleaned-longmemeval-s.longmemeval-cn
LongMemEval-CN 500题中文子集|识流公开结果
这是识流维护的 LongMemEval 500题中文翻译子集与逐题评测结果。500题结果快照版本为 2026.07.15;本存档修订版为 2026.07.18,仅完善许可、引用与跨平台存档元数据,results.jsonl 未改变。
该子集状态为 draft,不是 LongMemEval 官方发布的中文版本。源自 LongMemEval 的问题与参考答案继续遵循上游 MIT License;由识流新增且有权授权的中文译文、模型输出整理、评测结果、汇总数据与原创说明采用 CC BY 4.0。完整许可边界见 NOTICE.txt。
结果摘要
测试日期:2026-07-15
总题数:500
跳过:0
首轮通过:499/500(99.8%)
独立复判:唯一未通过题复判通过
复核后有效通过:500/500
exact-match guard:298题
DeepSeek deepseek-v4-flash 判分:202题… See the full description on the dataset page: https://huggingface.co/datasets/shiliu-memory/longmemeval-cn.longmemeval-s
LIXINYI33/longmemeval-s
This dataset repository contains JSON files compatible with LongMemEval-like experiments.
Uploaded files:
longmemeval_s_cleaned.json
Notes:
This repo was created/updated via a utility script and is intended for personal experiments.
JSON files can be loaded directly with the datasets library using the json builder, e.g.:
from datasets import load_dataset
# load a single file from the Hub
# replace <file.json> with your target file name
# requires… See the full description on the dataset page: https://huggingface.co/datasets/LIXINYI33/longmemeval-s.longmemeval-qwenreminisce-longmemeval-results
Reminisce LongMemEvalS Benchmark Results
Benchmark results for Reminisce, a cognitive science-inspired memory architecture for AI agents.
Benchmark
Dataset: LongMemEvalS (Wu et al., ICLR 2025) - 500 questions across 6 categories
Retrieval: Keyword-overlap ranking over episodic memory summaries (no vector search, no consolidation)
Judge: Qwen3 80B (local)
Models tested: Claude Haiku 4.5, Claude Sonnet 4.6, Claude Opus 4.6
Results
Model
Overall
Precision… See the full description on the dataset page: https://huggingface.co/datasets/myronkoch/reminisce-longmemeval-results.longmemeval-oracle-ko
LongMemEval oracle (한국어 번역)
LongMemEval (Wu et al., ICLR 2025) 의 oracle 분할 를 한국어로 기계 번역한 자료입니다. 번역 벤치마크와 한국어로 직접 작성한 벤치마크를 비교하는 연구를 위해 만들었습니다.
불러오기
from datasets import load_dataset
ds = load_dataset("NAMJOON/longmemeval-oracle-ko") # 한국어 + 원문
ds = load_dataset("NAMJOON/longmemeval-oracle-ko", "korean_only") # 한국어만
구성
500 문항, 증거 세션 948 개. 원본 구조를 유지합니다.
question_type: 6종 (temporal-reasoning, multi-session, knowledge-update… See the full description on the dataset page: https://huggingface.co/datasets/NAMJOON/longmemeval-oracle-ko.longmemeval-results
Feather DB — LongMemEval Benchmark Results
Feather DB v0.8.0 results on LongMemEval (ICLR 2025) — the long-term memory benchmark for chat assistants.
Results
Configuration
Variant
Score
Cost
Feather + GPT-4o
S
0.693
~$8
Feather + Gemini-2.5-Flash
S
0.657
~$2.40
Full-context GPT-4o (paper baseline)
S
0.640
—
Per-axis (S variant, GPT-4o)
Axis
Score
Information-extraction
0.942
Multi-session reasoning
0.606
Temporal reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Hawky-ai/longmemeval-results.MemCoT-longmemevalReMe_longmemeval_clean_s_v2
LongMemEval ReMe Cleaned-S
longmemeval_s_reme_cleaned.json is a corrected version of the LongMemEval
Cleaned-S dataset. It keeps the original questions and haystack sessions while
replacing the answer and supporting-session ground truth with the reviewed
values from final_groundtruth_cleaned_s.json.
The corrections address inaccurate answers and evidence sessions, including
cases where evidence occurred after the question time and therefore leaked
future information into the… See the full description on the dataset page: https://huggingface.co/datasets/agentscope-ai/ReMe_longmemeval_clean_s_v2.longmemeval-m-cleanedlongmemeval-v2This dataset replaces the original LongMemEval dataset. The main difference is that this version removes noisy history sessions that interfere with the answer correctness. More detailed session processing information can be found here.
LongMemEval-Slongmemeval-cleanedThis dataset replaces the original LongMemEval dataset. The main difference is that this version removes noisy history sessions that interfere with the answer correctness. More detailed session processing information can be found here.
longmemeval-pooled
LongMemEval Pooled — experience/question sets for memory & KV-cache (cartridge) research
Conversational experiences (contexts to hold in a window, or to distill into a compact
KV representation) paired with objectively-graded validation questions, derived from the
distractor sessions of LongMemEval (longmemeval_m,
cleaned release). Every answer is a single word/value from the source conversation — no LLM
judge needed.
The family is a 2×2×2 grid:
pooled3 / pooled4 — exactly 3… See the full description on the dataset page: https://huggingface.co/datasets/hhy13/longmemeval-pooled.LongMemEval-evallongmemeval-pooled-mcq
LongMemEval Pooled MCQ — context/question sets for KV-cache (cartridge) training
Conversational experiences (context to distill into a compact KV representation) paired with
exact-match 4-option MCQ validation questions, derived from the distractor sessions of
LongMemEval (longmemeval_m, cleaned release). No LLM judge
needed: answers are single tokens, graded by letter match.
config
records
contents
pooled3_experiences
88
conversations (~245k tok total), 3 questions… See the full description on the dataset page: https://huggingface.co/datasets/hhy13/longmemeval-pooled-mcq.LongMemEvallongmemeval⚠️ This dataset is deprecated. It is replaced by longmemeval-cleaned (https://huggingface.co/datasets/xiaowu0162/longmemeval-cleaned) which noisy history sessions that interfere with the answer correctness.
LongMemEval-S-eval
