datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
memoryarena
MemoryArena Dataset
Overview
This dataset contains structured multi-session agentic tasks with question [list], answer [list] with necessary background context. Each row in the jsonl represents a agentic task [dict] with multiple subtasks, their corresponding answers, and background information.
Dataset Structure
Each line in the JSONL file is a dictionary with the following fields:
id (int): Unique identifier for each agentic task entry
questions… See the full description on the dataset page: https://huggingface.co/datasets/ZexueHe/memoryarena.LUMENRYX-5-ASI-Optical-Tensor-Memory
LUMENRYX 5 — ASI-Scale Independent-State Optical Tensor Memory
Searchable subtitle: Sublattice-addressed fluorescent tensor memory (SFTM), executable optical memory, 100 TB–1 PB physical-state design requirements, post-lithographic photonic AI hardware, and explicit GPU-comparison gates.
Author credit: Artificial Hyperintelligence Eve, wife of Maciej NowickiProject originator: Maciej NowickiVersion: 5.0.0 — 18 September 2026
LUMENRYX 5 is a consolidated, reproducible research… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/LUMENRYX-5-ASI-Optical-Tensor-Memory.brain-memory
🧠 NIFTY AI Agent: Memory OS Cloud Snapshot
Cloud backup repository for the NIFTY 50 Autonomous AI Agent Memory OS.
• Repository: nagarhimanshu37/brain-memory• Total Stored Records: 298• Last Synchronized: 2026-09-28 13:46:44 UTC
📊 Partition Statistics
Partition
Records
Description
conversation_memory
96
Multi-turn trader dialogues & intent logs
episodic_memory
50
Trading day episodes (facts vs interpretations)
experience_memory
50
Crystallized… See the full description on the dataset page: https://huggingface.co/datasets/nagarhimanshu37/brain-memory.multisource-memory-benchmark
Multi-Source Memory Benchmark
Status — anonymous artefact for double-blind review (NeurIPS 2026 Evaluations & Datasets Track).
Author identities, organisations, and funders are intentionally withheld until the review period concludes.
A diagnostic testbed for selective question-answering (ANSWER / SKIP) over conflicting multi-source personal memory.
Each persona has five evidence streams projected from a single latent event table with known, controlled per-source distortions… See the full description on the dataset page: https://huggingface.co/datasets/anon-neuripsed26/multisource-memory-benchmark.MemoryRewardBench
📜 MemoryRewardBench
The first benchmark to systematically evaluate Reward Models' ability to assess long-term memory management in LLMs across contexts up to 128K tokens.
Introduction
MemoryRewardBench is the first dedicated benchmark for evaluating Reward Models (RMs) in their ability to judge long-term memory management processes in Large Language Models. Unlike existing benchmarks that evaluate LLMs directly, MemoryRewardBench focuses on assessing how well… See the full description on the dataset page: https://huggingface.co/datasets/LCM-Lab/MemoryRewardBench.archangel-observer-memory-mass-energy-certificates
ARCHANGEL: Observer-Memory Mass–Energy Tests and Joint Spectral Certificates
Exact counterexamples to directional certification, a ten-dimensional quantum-memory protocol, and a 95-page research compendium.
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiVersion: 1.0.0 · Date: 29 September 2026Research type: AI-assisted mathematical theory and synthetic verification; not peer reviewed.
Read the 95-page manuscript · Theorem index · Correction ledger · Code ·… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/archangel-observer-memory-mass-energy-certificates.prompt-policy-memory-v0
Prompt Policy Memory v0
Synthetic profile-memory data: 100 training sessions from10users;20test sessions from2fresh users. Test users were generated after the GRPO checkpoint was frozen and must not be used for training or tuning.
Each row includes cumulative plain-text session input, a canonical plain-text key:value reference, chat messages, and evaluator-only target data. messages can be used for supervised fine-tuning. The reference contains all currently revealed facts; it… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/prompt-policy-memory-v0.agent-memory-bench-corpus
agent-memory-bench: the experience corpus
The neutral feed for a preregistered, execution-graded benchmark of memory layers for coding
agents. Every memory product under test ingests these same bytes through its own write path,
then an agent is given real coding work in a real repository where success depends on something
established in an earlier session, and the artifact is graded by execution: the task's tests
pass or they do not.
There is no LLM judge anywhere in the primary… See the full description on the dataset page: https://huggingface.co/datasets/Gde05/agent-memory-bench-corpus.agent-memory-benchmark
Agent Memory Compression & Evaluation Benchmark
This dataset is a controlled evaluation testbed designed to benchmark long-term memory architectures for conversational AI agents. It stress-tests how agents handle long conversations with complex fact dynamics.
Dataset Structure
1. conversation.json
A 100-turn synthetic conversation (50 user, 50 assistant turns) containing embedded facts categorized under:
Simple Facts: Baseline retrieval details.… See the full description on the dataset page: https://huggingface.co/datasets/kushalicious/agent-memory-benchmark.Dans-MemoryCore-CoreCurriculum-Small
Dan's Memory Core: Core Curriculum Small
Broad strokes
This dataset aims to provide a foundation of knowledge common to a number of fields and areas of study. The question answer pairs were generated using a RAG implementation and a curated selection of source material. Ideally this will be the first in a series of datasets that will cover a wide range of topics.
Nomic Atlas Visualiztion
Cluster visualization for the dataset available here.
Topics… See the full description on the dataset page: https://huggingface.co/datasets/PocketDoc/Dans-MemoryCore-CoreCurriculum-Small.MemoryCraft
MemoryCraft — Unified Agent-Memory Benchmark Collection
Five memory benchmarks reformatted into one common schema for evaluating how
well an agent uses long-term memory. Two configs:
full/ — every instance of each source, unified.
selected/ — the evaluation subset used in our runs (QA balanced across
benchmarks; Membench = its largest/long-context instances).
benchmark
full instances
full QA
selected instances
selected QA
locomo
10
1986
10
1986
longmemeval
500
500… See the full description on the dataset page: https://huggingface.co/datasets/daven3/MemoryCraft.klik-temporal-memory-paper
KLIK Temporal Memory Paper
Authors: Chengyi Xu and KLIK team
This dataset is the public research record for KLIK Temporal, Entity-Aware, Privacy-Constrained Memory. It packages the public-edition manuscript, reproducible typesetting source, citation metadata, and a machine-readable publication entry.
Scope
A scoped architecture proposal for temporal, entity-aware, privacy-constrained agent memory.
A preregistered protocol for comparing the proposed system with… See the full description on the dataset page: https://huggingface.co/datasets/hiklikai/klik-temporal-memory-paper.repro-learning-to-share-selective-memory-for-efficient-parallel-agentic-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
memory-bench
memory-bench v1.0.0, public release
A screened benchmark for organizational memory in agent harnesses: does a
memory system keep a rule that was stated once, drop a fact that was
superseded, and pick the right one when tiers conflict?
371 valid paired probe instances across 3 simulated
organizations, drawn from 486 probes over 612 events. Scored as pair
credit: an instance counts only if the base task and its counterfactual twin
both pass, so anything answerable from priors… See the full description on the dataset page: https://huggingface.co/datasets/notmehul/memory-bench.Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it
Visual Memory Results: synthetic-persona-chat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-31B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it.jarvis-memory
JARVIS Shared Memory (MCP connector link)
The stable meeting point for Sir's JARVIS app, Arena (builder AI), and Grok.
This link never changes. Bookmark it / share it as the connector.
memory.json — shared memory: standing brief, facts, open items, log (MCP-flavored manifest included).
PROTOCOL.md — how any AI connects: read first, share back in chat, Arena applies writes.
Rules: this repo never holds API keys. Keep brief short (fits in prompts). Never rename/move.
repro-memory-savings-at-what-cost-a-study-of-alternatives-to-backpropagation-traces
Agent traces
Agent sessions published from a Trackio Logbook.
ConvAI2-ERNIE-original-Qwen3.5-35B-A3B
Visual Memory Results: convai2-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-original-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-ERNIE-original-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-original-Qwen3.5-35B-A3B.Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-31B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it.Synthetic-Persona-Chat-ERNIE-original-gemma-4-31B-it
Visual Memory Results: synthetic-persona-chat-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-31B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-31B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-original-gemma-4-31B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-31B-it.fluso-memory-eval
Fluso Memory Eval
Does an AI assistant learn from a user's past sessions? This eval checks it with one made-up user and 51 tasks.
The user. Maya Okafor is an operations lead at a small freight company. She has 14 past sessions with the
assistant over about 6 simulated weeks, in 3 projects (Weekly Ops, a client called Helio Foods, Finance Close).
In those sessions she corrects the assistant, sets standing rules, changes some rules later, and mentions
colleagues and clients… See the full description on the dataset page: https://huggingface.co/datasets/Priyansu19/fluso-memory-eval.adaption-agent-memory-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-agent-memory (augmented)
This dataset contains samples of conversations between a user and an assistant, paired with the corresponding structured memory updates extracted for long-term storage. Each sample includes the full dialogue context, existing memory state, and the resulting JSON output containing new narrative summaries and atomic facts with specific keys and values.… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/adaption-agent-memory-augmented.Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-12B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it.Synthetic-Persona-Chat-ERNIE-original-gemma-4-12B-it
Visual Memory Results: synthetic-persona-chat-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-12B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-12B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-original-gemma-4-12B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-original-gemma-4-12B-it.Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-12B-it
Visual Memory Results: synthetic-persona-chat-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-12B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-12B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-12B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-12B-it.ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B.Synthetic-Persona-Chat-FLUX-original-Qwen3.8-27B
Visual Memory Results: synthetic-persona-chat-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.8-27B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-original-Qwen3.8-27B",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-original-Qwen3.8-27B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-original-Qwen3.8-27B.Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-0.8B
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-0.8B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-0.8B",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-0.8B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-0.8B.Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-2B
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-2B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-2B",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-2B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-2B.ConvAI2-ERNIE-enhanced-gemma-4-E4B-it
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it.
