compaction
agent-memory-compaction-trajectories
Agent Memory Compaction Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/agent-memory-compaction-trajectories.agenttrove-glm53-compactions
AgentTrove compactions
10,000 summaries generated by GLM-5.3 from GLM-4.6 traces in
AgentTrove, using
OpenCode's compaction prompt. This dataset stores summaries and source references,
not the original traces.
Marin SFT token count
Marin's SFT representation contains 188,537,557 tokens after chat normalization,
rendering, and tokenization with its 2026-09-18 tokenizer snapshot. The count includes
AgentTrove histories restored into the compaction prompts.… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/agenttrove-glm53-compactions.drctx-compaction-openresearcher-mk1context-compaction-poc
Context Compaction PoC Dataset
Training data for a context compaction model — a model that decides which lines to KEEP vs DROP from coding agent context (file reads, grep results, test output, etc.).
Every surviving line stays verbatim. No summarization, no rewriting, zero hallucination risk. Dropped lines become (filtered N lines) markers.
Why context compaction?
Coding agents (Claude Code, Codex, SWE-agent) accumulate massive context during long sessions — 70%+… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/context-compaction-poc.openresearcher-compaction-mk2-nc1tmax-compaction-effort-nc0.3
