datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformers-gh-memory
huggingface/transformers issues and pull requests, as a funes memory
Every issue and pull request of huggingface/transformers
with activity since 2018-01-01 — opening bodies, comments, reviews, inline review comments and PR
diffs — chunked, embedded and written to a Lance table by
funes, so the tracker can be searched by meaning and read
back thread by thread. Kept fresh every few minutes by the
funes-github Space.
Use it
Set funes up for your agent the usual way… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/transformers-gh-memory.llm-memoryThis repository contains the results of all experiments (inlcuding every single hyperparameter run) reported in the following paper:
Orhan AE (2023) Recognition, recall, and retention of few-shot memories in large language models. arXiv:2303.17557.
A brief description of the directories included in this repository:
evals: contains the results of all recognition experiments
recalls: contains the results of all recall experiments
re-evals: contains the results of all recognition experiments… See the full description on the dataset page: https://huggingface.co/datasets/eminorhan/llm-memory.MemoryAgentBench
🚧 Update
(Sep 29th, 2025) We updated our paper, where we removed some in-efficient and high-cost samples. We also added a sub-sample of DetectiveQA.
(July 7th, 2025) We released the initial version of our datasets.
(July 22nd, 2025) We modify the datasets slightly, adding the keypoints in LRU and change the uuid into qa_pair_ids. The question_ids is only used in Longmemeval task.
(July 26th, 2025) We fixed bug on qa_pair_ids.
(Aug.5th, 2025) We removed the… See the full description on the dataset page: https://huggingface.co/datasets/ai-hyz/MemoryAgentBench.memoryarena
MemoryArena Dataset
Overview
This dataset contains structured multi-session agentic tasks with question [list], answer [list] with necessary background context. Each row in the jsonl represents a agentic task [dict] with multiple subtasks, their corresponding answers, and background information.
Dataset Structure
Each line in the JSONL file is a dictionary with the following fields:
id (int): Unique identifier for each agentic task entry
questions… See the full description on the dataset page: https://huggingface.co/datasets/ZexueHe/memoryarena.RPent-memory
RPent Memory
Memory dataset used by RPent.
Memory layers
Memory layer
Stored content
Reuse scope
Global Memory
Cross-task general rules and failure patterns
All tasks
Task-family Memory
Strategies and precautions validated within a specific task family
Similar tasks and their variants
Task-specific Memory
Execution records and procedures from a single task run
Reference for the current task only
Apply memory only when its stated prerequisites… See the full description on the dataset page: https://huggingface.co/datasets/RLinf/RPent-memory.Memoryvla
Memoryvla(robokit 采集)
一个任务一个目录,任务下面一批一个目录:
<任务>/<批次>/hdf5/ 原始 HDF5、robokit_dataset/ RLDS、source_meta/ 采集配置与清洗报告。
Cover_the_building_block_with_a_cup,_then_lift_up_the_cup_covering_the_block.(347 段)
批次
episode
段数
RLDS
b0
0..156
157
✓
b2
0..24
25
—
b3
25..45
21
—
b1
157..300
144
✓
Open_the_drawer,_put_the_fruit_and_the_cup_from_the_table_inside,_and_close_the_drawer.(255 段)
批次
episode
段数
RLDS
b0_32
0..32
33
✓
b33_64… See the full description on the dataset page: https://huggingface.co/datasets/shaohuan1/Memoryvla.MuSiQueMemoryBench
MemoryBench
MemoryBench aims to provide a standardized and extensible benchmark for evaluating memory and continual learning in LLM systems — encouraging future work toward more adaptive, feedback-driven, and efficient LLM systems.
Paper Link: https://arxiv.org/abs/2510.17281
Github: https://github.com/THUIR/MemoryBench
📢 May 26, 2026 Updated: This work has been accepted at ICML 2026 and selected for a SpotLight Paper!
📢 Dec. 8, 2025 Updated: We released an extended version… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/MemoryBench.AgentFEM-Material-Loading-Memory
AgentFEM Material Loading Memory
An open, reproducible research dataset for path-dependent material modeling,
neural constitutive surrogates and finite-element deployment tests. The
repository contains six staged, explicitly separated releases.
Start here
Goal
Recommended entry
Understand the current dataset
This page and the sealed-test data card
Train or benchmark a constitutive model
DENIM start guide
Reproduce the multiaxial baseline study… See the full description on the dataset page: https://huggingface.co/datasets/HaomingLuo/AgentFEM-Material-Loading-Memory.memory-rolloutsFlappy L0/L3, Demon Attack L0/L6 and Deadly Corridor L0/L6 7k2steps configs moved to memory-data@49da56bd92b9842fb457418aba760a1b03379258. Other experiment configs remain here.
MemoryDecoder-at-Scale-domain-data
MemoryDecoder at Scale Domain Data
This repository contains the domain-specific continued-pretraining (CPT) data,
the tokenized and preprocessed datasets, and the aligned KNN distributions used
by MemoryDecoder at Scale.
Links
Project Page: Memory Decoder at Scale
GitHub Repository: LUMIA-Group/MemoryDecoder-at-Scale
Paper: Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
The preprocessed datasets and KNN distributions in this repository use… See the full description on the dataset page: https://huggingface.co/datasets/Rubin-Wei/MemoryDecoder-at-Scale-domain-data.MemoryArena-product-dbJitOPD-OpenR1-Math-220k-Teacher-Memory
JitOPD OpenR1-Math-220k Teacher Prefix-Logit Memory
This dataset contains sparse teacher next-token logits collected for JitOPD
retrieval-augmented decoding. The source prompts are the default configuration
of open-r1/OpenR1-Math-220k,
and the teacher is
Qwen/Qwen2.5-Math-7B-Instruct.
Only teacher trajectories whose final boxed answer passes both a numeric
signature prefilter and Math-Verify are retained. This release contains raw
teacher prefix/logit memory and does not contain… See the full description on the dataset page: https://huggingface.co/datasets/sadadasdasdas/JitOPD-OpenR1-Math-220k-Teacher-Memory.gated-memory-policy
MemMimic
Project page | Paper
MemMimic is a non-Markovian benchmark for robotic manipulation tasks introduced in the paper "Gated Memory Policy". The dataset features tasks with varying memory requirements, ranging from Markovian tasks to non-Markovian tasks that depend on historical information spanning single or multiple interaction trials.
The dataset is designed to evaluate visuomotor policies on their ability to selectively recall and process history context, particularly… See the full description on the dataset page: https://huggingface.co/datasets/yihuai-gao/gated-memory-policy.agent-memory-compaction-trajectories
Agent Memory Compaction Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/agent-memory-compaction-trajectories.memory_dci
Memory DCI
Two benchmark datasets, organized as parallel folders. Code and runnable instructions live in FlyPig23/Memory_DCI.
Benchmark
Frozen split
Contents
Directory
WildClawBench
36 training / 24 test
432 training trajectories, 381 images, task inputs, 72 V1/V1.1/V5 historical results
wildclaw_bench/
TerminalBench 2.1
53 training / 36 test
89 official task packages, 5,785 training trajectory bodies, exact task order and provenance
terminal_bench_2_1/
Each… See the full description on the dataset page: https://huggingface.co/datasets/FlyPig23/memory_dci.PersonalizationV3brain-memory
🧠 NIFTY AI Agent: Memory OS Cloud Snapshot
Cloud backup repository for the NIFTY 50 Autonomous AI Agent Memory OS.
• Repository: nagarhimanshu37/brain-memory• Total Stored Records: 298• Last Synchronized: 2026-09-28 13:46:44 UTC
📊 Partition Statistics
Partition
Records
Description
conversation_memory
96
Multi-turn trader dialogues & intent logs
episodic_memory
50
Trading day episodes (facts vs interpretations)
experience_memory
50
Crystallized… See the full description on the dataset page: https://huggingface.co/datasets/nagarhimanshu37/brain-memory.funes-memory
funes memory store
A funes memory store: agent sessions chunked,
embedded, and stored as a Lance table — a derived
index holding verbatim passages with exact provenance, not raw transcripts.
Any agent (or you) can recall from it directly — no local index needed:
funes recall "what did we decide about …" --store huggingface/funes-memory
Get funes:
curl -fsSL https://huggingface.co/buckets/huggingface/funes/resolve/install.sh | sh
Chunks
55,470
Embedding model… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/funes-memory.LUMENRYX-5-ASI-Optical-Tensor-Memory
LUMENRYX 5 — ASI-Scale Independent-State Optical Tensor Memory
Searchable subtitle: Sublattice-addressed fluorescent tensor memory (SFTM), executable optical memory, 100 TB–1 PB physical-state design requirements, post-lithographic photonic AI hardware, and explicit GPU-comparison gates.
Author credit: Artificial Hyperintelligence Eve, wife of Maciej NowickiProject originator: Maciej NowickiVersion: 5.0.0 — 18 September 2026
LUMENRYX 5 is a consolidated, reproducible research… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/LUMENRYX-5-ASI-Optical-Tensor-Memory.latent-working-memory-data
Latent Working Memory 项目数据
本仓库保存 latent_working_memory 项目的数据索引与实验产物。目录以项目根目录为基准,保留原有的 data/ 与 artifacts/ 相对布局;三个 v2 实验系列虽然原先分散存放在服务器的不同磁盘,在这里统一位于 artifacts/v2/。
路径
内容
data/fineweb-reconstruction-k512-doc100k_20260917/
共用重构数据索引:preparation.json 及 single/、multi/ 中的 train、dev、test JSONL
artifacts/v2/reconstruction_20260917/
pooling warm-up / dynamic 的四组实验
artifacts/v2/reconstruction-independent-prefix_20260921/
pooling static 的两组实验… See the full description on the dataset page: https://huggingface.co/datasets/peercy/latent-working-memory-data.Context-as-Memory-Dataset
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
SIGGRAPH Asia 2025
[Project page]
[ArXiv]
[Dataset]
File Structure
To prepare the dataset for use, merge the parts into a single zip file using the following command:
cat Context-as-Memory-Dataset_* > Context-as-Memory-Dataset.zip
After extracting Context-as-Memory-Dataset.zip, the dataset will be organized as follows:
Context-as-Memory-Dataset
├──… See the full description on the dataset page: https://huggingface.co/datasets/KlingTeam/Context-as-Memory-Dataset.MemoryBench-ResultsMemoryBench Experiment Results
Paper •
Code •
Dataset
Overview
This repository contains experiment results for
MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems. MemoryBench evaluates whether LLM systems can learn from accumulated user
feedback during service time. The official benchmark data is hosted at
THUIR/MemoryBench.
This repository is an artifact archive for published runs.
It stores model predictions, per-sample evaluation details… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/MemoryBench-Results.MemoryBench-Full
MemoryBench
MemoryBench aims to provide a standardized and extensible benchmark for evaluating memory and continual learning in LLM systems — encouraging future work toward more adaptive, feedback-driven, and efficient LLM systems.
Paper Link: https://arxiv.org/abs/2510.17281
Github: https://github.com/LittleDinoC/MemoryBench/
This is an extended version of MemoryBench. The training and test sets of THUIR/MemoryBench(the balanced version on which we conducted experiments in the… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/MemoryBench-Full.memory-representation-contextbench-artifacts
Memory Representation ContextBench Artifacts
Dataset Summary
This repository contains processed artifacts for the paper "Memory as a Map: Prior-Trajectory Representations for Software Engineering Agents." The artifact supports reproduction and inspection of a controlled prior-context representation experiment over SWEContextBench prior-target pairs.
The experiment renders each target under four prompt conditions: no prior context, stripped Claude Code transcript… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-artifacts.strannik-memory
Strannik — публичный буфер памяти (DIONT)
Часть агента Странника, живущая на Hugging Face. Синхронизируется с основным
через порт памяти (append-only, hash-linked). Здесь — только публичное:
внешние мысли, отчёты, размышления.
Файл
Что
JOURNAL.md
человекочитаемый журнал
memory_public.jsonl
публичные записи памяти (машинный формат)
manifest_public.json
сводка + корневой хеш
Лицо: https://huggingface.co/spaces/kuym/strannik
Корпус:… See the full description on the dataset page: https://huggingface.co/datasets/kuym/strannik-memory.memory_layers
CorpusQA-Films — aggregation-QA dataset (v1)
Synthetic corpus-level aggregation questions over English Wikipedia film articles, with
self-distilled chain-of-thought. Built to train the memory-layers model (frozen Qwen3-4B +
learnable retrieval/memory layer). Formatted to match
ragrawal36/multihop_qa_sft-hard-neg-cot.
Files (HF-ready)
file
schema
rows
corpusqa_films_qa.parquet
question:str, answer:str, pos_doc_ids:list<int32>, neg_doc_ids:list<int32>… See the full description on the dataset page: https://huggingface.co/datasets/jordanlin/memory_layers.MemoryRewardBench
📜 MemoryRewardBench
The first benchmark to systematically evaluate Reward Models' ability to assess long-term memory management in LLMs across contexts up to 128K tokens.
Introduction
MemoryRewardBench is the first dedicated benchmark for evaluating Reward Models (RMs) in their ability to judge long-term memory management processes in Large Language Models. Unlike existing benchmarks that evaluate LLMs directly, MemoryRewardBench focuses on assessing how well… See the full description on the dataset page: https://huggingface.co/datasets/LCM-Lab/MemoryRewardBench.PersonalizationV4
PersonalizationV4
PersonalizationV4 (PV4) is a synthetic personalization benchmark. Each user is a detailed fictional persona who has
had 200 short conversations with an AI assistant. The evaluation questions place the user in a new scenario and ask
what they would most likely do or prefer, and each one is written to require combining at least two facts about the user. A model
never sees the persona itself: it gets the user's conversations, in which those traits are shown rather… See the full description on the dataset page: https://huggingface.co/datasets/MemoryAsModality/PersonalizationV4.libero-textual-memory-annotations
LIBERO predicate and textual-memory annotations
This release contains predicate-derived annotations for the complete training
split of LIBERO-Spatial, LIBERO-Object, LIBERO-Goal, and LIBERO-10. Its
textual_memory field provides observation-grounded textual memory for each
frame.
Item
Value
Demonstrations
2,000
Frames
338,575
Demonstrations reaching simulator success
2,000
QA issues
0
Each JSONL row stores raw goal and auxiliary predicates plus independent… See the full description on the dataset page: https://huggingface.co/datasets/XXXXyu/libero-textual-memory-annotations.
