Team Ai
Datasetpublic

cozzyde/long-horizon-memorization

ComposeCL Datasets This repository contains the three 100-task memorization datasets from the paper Continual Learning Mechanisms Compose for Long-Horizon Memorization (arXiv: 2609.06986). Project page: https://compose-cl.github.io/ Code: https://github.com/cozheyuanzhangde/compose-cl Datasets Dataset Path Tasks Items/task Construction Symbol-QA data/synthetic_qa/symbol_qa 100 100 random six-character keys mapped to random four-character values… See the full description on the dataset page: https://huggingface.co/datasets/cozzyde/long-horizon-memorization.

sourceHugging Facemitupdated 24d agoView on Hugging Face
3likes198downloads
Dataset Card

ComposeCL Datasets

This repository contains the three 100-task memorization datasets from the paper Continual Learning Mechanisms Compose for Long-Horizon Memorization (arXiv: 2609.06986).

  • —Project page: https://compose-cl.github.io/
  • —Code: https://github.com/cozheyuanzhangde/compose-cl

Datasets

DatasetPathTasksItems/taskConstruction
Symbol-QAdata/synthetic_qa/symbol_qa100100random six-character keys mapped to random four-character values
LLM-QAdata/synthetic_qa/llm_qa100100LLM-generated fictional facts, with task order shuffled using seed 0
Real-QAdata/real_qa/qwen3_4b_base10050ten public QA sources, filtered against Qwen3-4B-Base and globally remixed

Training and evaluation use the same associations because the target quantity is memorization retention rather than held-out generalization. Answers are rendered as \boxed{answer} at load time. The stored JSON remains plain text.

The included Real-QA questions were filtered against Qwen/Qwen3-4B-Base and should be paired with that backbone when reproducing the paper.