longmemeval
longmemeval-cleanedThis dataset replaces the original LongMemEval dataset. The main difference is that this version removes noisy history sessions that interfere with the answer correctness. More detailed session processing information can be found here.
longmemeval⚠️ This dataset is deprecated. It is replaced by longmemeval-cleaned (https://huggingface.co/datasets/xiaowu0162/longmemeval-cleaned) which noisy history sessions that interfere with the answer correctness.
longmemeval-v2
LongMemEval-V2 Data
Project Page | Paper | GitHub
LongMemEval-V2 (LME-V2) is an evaluation benchmark for long-term memory in web and enterprise agents. It contains 451 manually curated questions and 1,870 task trajectories drawn from customized WebArena-style and ServiceNow-style environments.
The benchmark evaluates five core memory abilities:
Static state recall: remembers important landmarks and page layouts.
Dynamic state tracking: understands how states change over time.… See the full description on the dataset page: https://huggingface.co/datasets/xiaowu0162/longmemeval-v2.funes-xiaowu0162-longmemeval-cleaned-s
Funes recall store — LongMemEval_s cleaned corpus
A funes recall store built by indexing the
longmemeval_s_cleaned.json haystack of
xiaowu0162/longmemeval-cleaned
(LongMemEval, arXiv:2410.10813) — every unique
chat session across all 500 questions' haystacks, in one corpus-wide store.
What this is
This is not a raw trace dataset — it is a pre-built funes index: the source
sessions chunked into content blocks and embedded, stored as a
Lance table (chunks.lance).… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-xiaowu0162-longmemeval-cleaned-s.longmemeval-s-cleanedLongMemEvalDoc
