datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MemoryBench
MemoryBench
MemoryBench aims to provide a standardized and extensible benchmark for evaluating memory and continual learning in LLM systems — encouraging future work toward more adaptive, feedback-driven, and efficient LLM systems.
Paper Link: https://arxiv.org/abs/2510.17281
Github: https://github.com/THUIR/MemoryBench
📢 May 26, 2026 Updated: This work has been accepted at ICML 2026 and selected for a SpotLight Paper!
📢 Dec. 8, 2025 Updated: We released an extended version… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/MemoryBench.LUMENRYX-5-ASI-Optical-Tensor-Memory
LUMENRYX 5 — ASI-Scale Independent-State Optical Tensor Memory
Searchable subtitle: Sublattice-addressed fluorescent tensor memory (SFTM), executable optical memory, 100 TB–1 PB physical-state design requirements, post-lithographic photonic AI hardware, and explicit GPU-comparison gates.
Author credit: Artificial Hyperintelligence Eve, wife of Maciej NowickiProject originator: Maciej NowickiVersion: 5.0.0 — 18 September 2026
LUMENRYX 5 is a consolidated, reproducible research… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/LUMENRYX-5-ASI-Optical-Tensor-Memory.PersonalizationV3memory_layers
CorpusQA-Films — aggregation-QA dataset (v1)
Synthetic corpus-level aggregation questions over English Wikipedia film articles, with
self-distilled chain-of-thought. Built to train the memory-layers model (frozen Qwen3-4B +
learnable retrieval/memory layer). Formatted to match
ragrawal36/multihop_qa_sft-hard-neg-cot.
Files (HF-ready)
file
schema
rows
corpusqa_films_qa.parquet
question:str, answer:str, pos_doc_ids:list<int32>, neg_doc_ids:list<int32>… See the full description on the dataset page: https://huggingface.co/datasets/jordanlin/memory_layers.MemoryBench-Full
MemoryBench
MemoryBench aims to provide a standardized and extensible benchmark for evaluating memory and continual learning in LLM systems — encouraging future work toward more adaptive, feedback-driven, and efficient LLM systems.
Paper Link: https://arxiv.org/abs/2510.17281
Github: https://github.com/LittleDinoC/MemoryBench/
This is an extended version of MemoryBench. The training and test sets of THUIR/MemoryBench(the balanced version on which we conducted experiments in the… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/MemoryBench-Full.memory-integrity-continuity
AtMem memory integrity and continuity evidence
This dataset accompanies Beyond Recall Accuracy: Evaluating Integrity and
Crash Continuity in Persistent Memory for Tool-Using Language Agents by
Javad Taghia, AtMem.Ai Lab.
Open the interactive results Space
· Read the PDF
· AtMem source
· External benchmark submission
What is measured
The publication separates two questions:
Memory integrity: can untrusted observations, summaries, repetition, or
prior agent outputs… See the full description on the dataset page: https://huggingface.co/datasets/atmem/memory-integrity-continuity.memory-representation-contextbench-artifacts
Memory Representation ContextBench Artifacts
Dataset Summary
This repository contains processed artifacts for the paper "Memory as a Map: Prior-Trajectory Representations for Software Engineering Agents." The artifact supports reproduction and inspection of a controlled prior-context representation experiment over SWEContextBench prior-target pairs.
The experiment renders each target under four prompt conditions: no prior context, stripped Claude Code transcript… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-artifacts.MemoryMatters_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 50,
"total_frames": 14325,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wayrise/MemoryMatters_Dataset.helium_memory
Try gpt-oss ·
Guides ·
Model card ·
OpenAI blog
Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.
We’re releasing two flavors of these open models:
gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters)gpt-oss-20b — for lower latency, and local or… See the full description on the dataset page: https://huggingface.co/datasets/Fred808/helium_memory.agent-memory-resilience-benchmark
Agent Memory Resilience & Poisoning Benchmark
Dataset Summary
This benchmark dataset evaluates resilience, negative transfer, and memory poisoning mitigation in autonomous LLM agent architectures (such as LangGraph, AutoGen, and CrewAI).
When autonomous agents record distilled self-reflections after attempting tasks, external stochastic failures or subtle API deprecations often cause agents to commit defective strategies into episodic memory. Under standard… See the full description on the dataset page: https://huggingface.co/datasets/sumitaidev/agent-memory-resilience-benchmark.privacy-aware-memory-benchmark
Privacy-Aware Memory Benchmark
Paper | Code
This repository provides synthetic, multi-turn, privacy-aware conversation histories used in our SP-Mem paper. The conversations contain both private information and non-private preferences across education, finance, medical, and mental domains.
Data
The conversation histories are stored as JSON files, with one file per synthetic user, organized by domain.
Domain
Users
Dialogue sessions
Education
250
5,250… See the full description on the dataset page: https://huggingface.co/datasets/wwj95/privacy-aware-memory-benchmark.memory-representation-contextbench-traces
Memory Representation ContextBench Raw Traces
This optional artifact contains raw Claude Code prior JSONL traces discovered for the ContextBench prompt set. It includes 96 trace manifest rows and 42722114 bytes of copied JSONL content.
OpenHands target-run JSONL traces were not present in the discovered source folders, so traces/openhands_runs/ is present as an empty directory structure and the absence is recorded in manifests/validation_summary.json.
Checksums are in… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-contextbench-traces.prompt-policy-memory-v0
Prompt Policy Memory v0
Synthetic profile-memory data: 100 training sessions from10users;20test sessions from2fresh users. Test users were generated after the GRPO checkpoint was frozen and must not be used for training or tuning.
Each row includes cumulative plain-text session input, a canonical plain-text key:value reference, chat messages, and evaluator-only target data. messages can be used for supervised fine-tuning. The reference contains all currently revealed facts; it… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/prompt-policy-memory-v0.PersonaMem-v2memory-representation-nebius-openhands-adp-v0.1
Nebius OpenHands Trajectories Converted to ADP v0.1
This is an auxiliary converter-validation artifact. It is not part of the 95-target SWEContextBench evaluation.
It converts public Nebius OpenHands trajectories to a canonical ADP-like JSON representation. The release contains 67074 converted rows across 17 parquet shard(s), covering 1823 repositories. Resolved labels: 32161 resolved and 34913 unresolved.
The artifact should be cited together with the original Nebius… See the full description on the dataset page: https://huggingface.co/datasets/shshwtsuthar/memory-representation-nebius-openhands-adp-v0.1.LongLaMP-2repro-learning-to-share-selective-memory-for-efficient-parallel-agentic-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
MemoryAgentBenchSingleRowThe same data as the original but put in single row format rather than multiple questions per line. Please see the original for questions, details, and the citation.
Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it
Visual Memory Results: synthetic-persona-chat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-31B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it.vln-trajectory-memory-stage2
VLN Trajectory-Memory — Stage 2 (projector alignment)
Text-only question answering where the only source of truth is a robot's action history.
Each record gives a navigation trajectory as a list of primitive actions and asks something
that can only be answered by tracking where those actions lead: how far from the start, which
way the robot faces, what happened in the last quarter of the route.
It was built to measure whether a frozen vision-language model (Qwen3-VL-2B) can read… See the full description on the dataset page: https://huggingface.co/datasets/Rithvik762/vln-trajectory-memory-stage2.repro-memory-savings-at-what-cost-a-study-of-alternatives-to-backpropagation-traces
Agent traces
Agent sessions published from a Trackio Logbook.
ConvAI2-ERNIE-original-Qwen3.5-35B-A3B
Visual Memory Results: convai2-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-original-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-ERNIE-original-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-original-Qwen3.5-35B-A3B.Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-31B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it.swebench-verified-kimi-k2p6-traces
SWE-bench Verified Kimi K2.6 Reasoning Traces
This dataset contains reasoning traces generated on princeton-nlp/SWE-bench_Verified using fireworks_ai/kimi-k2p6-high with a mini-swe-agent based harness. It is intended for research and distillation of software-engineering agents.
The repository is published with three configs because each table has a different schema:
raw_trajectories: one row per SWE-bench instance with the patch, sanitized result JSON, full trajectory JSON, message… See the full description on the dataset page: https://huggingface.co/datasets/MemoryAsModality/swebench-verified-kimi-k2p6-traces.Synthetic-Persona-Chat-ERNIE-original-gemma-4-31B-it
Visual Memory Results: synthetic-persona-chat-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-31B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-31B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-original-gemma-4-31B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-31B-it.eval_record-memory2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 4,
"total_frames": 7141,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jakacikac/eval_record-memory2.Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-12B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it.Synthetic-Persona-Chat-ERNIE-original-gemma-4-12B-it
Visual Memory Results: synthetic-persona-chat-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-12B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-12B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-original-gemma-4-12B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-12B-it.super_poulain_subtasks_memoryThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 50,
"total_frames": 32650,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/super_poulain_subtasks_memory.Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-12B-it
Visual Memory Results: synthetic-persona-chat-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-12B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-12B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-12B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-12B-it.
