datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-memory-compaction-trajectories
Agent Memory Compaction Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/agent-memory-compaction-trajectories.agenttrove-glm53-compactions
AgentTrove compactions
10,000 summaries generated by GLM-5.3 from GLM-4.6 traces in
AgentTrove, using
OpenCode's compaction prompt. This dataset stores summaries and source references,
not the original traces.
Marin SFT token count
Marin's SFT representation contains 188,537,557 tokens after chat normalization,
rendering, and tokenization with its 2026-09-18 tokenizer snapshot. The count includes
AgentTrove histories restored into the compaction prompts.… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/agenttrove-glm53-compactions.drctx-compaction-openresearcher-mk1context-compaction-poc
Context Compaction PoC Dataset
Training data for a context compaction model — a model that decides which lines to KEEP vs DROP from coding agent context (file reads, grep results, test output, etc.).
Every surviving line stays verbatim. No summarization, no rewriting, zero hallucination risk. Dropped lines become (filtered N lines) markers.
Why context compaction?
Coding agents (Claude Code, Codex, SWE-agent) accumulate massive context during long sessions — 70%+… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/context-compaction-poc.openresearcher-compaction-mk2-nc1tmax-compaction-effort-nc0.3cligym-compaction-effort-nc0.3tmax-compaction-mk2-nc1drctx-compaction-tmax-mk1compactionbench-lme-text-subset
CompactionBench LME Text Subset
20 questions from LongMemEval-V2 filtered for text-answerable content.
Each task has ~160k tokens of web agent trajectory context (thoughts, actions, page states).
All answers are confirmed present in the text context.
Used for testing context compaction in long-running agents.
Format
JSONL with one task per line. Each line is a CompactionBench TaskRow with:
task_id
context (~160k tokens)
question
gold_answer
metadata (domain… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/compactionbench-lme-text-subset.drctx-compaction-cligym-mk1cligym-compaction-mk2-nc1roller-compaction-ribbon-density
Roller Compaction: Ribbon Density vs. Process Parameters (Synthetic)
Version: 1.0
Publisher: Innovative Process Applications (IPA)
License: Creative Commons Attribution 4.0 International (CC BY 4.0)
Contact: Crestwood, IL, USA
⚠️ This dataset is 100% synthetic and intended for educational use only.
It was generated from a published physical model (Johanson rolling theory + Heckel densification) — not measured on any real equipment, customer, or production batch. Do not use it for… See the full description on the dataset page: https://huggingface.co/datasets/Innovative-Process-Applications/roller-compaction-ribbon-density.k3-dflash-sequence-compaction-validation-7cf713b
K3 DFlash Sequence-Compaction GB300 Validation
This repository is the evidence bundle for an independent correctness review of the
TensorRT-LLM K3 DFlash sequence-compaction patch at commit
7cf713b08b6892aae44a12d41b1a67d029d4b234.
The implementation was evaluated against the design contract at
f7789542915749fc9e6cd9b165b3a271cbe92184.
Bottom line
No patch correctness defect was found within the implemented and supported envelope. The patch's… See the full description on the dataset page: https://huggingface.co/datasets/srivy-together/k3-dflash-sequence-compaction-validation-7cf713b.Dry-Granulation_Roll-Compaction_and_Milling
Dry Granulation: Multi-Material Roll Compaction & Milling (Synthetic)
Version: 1.0
Publisher: Innovative Process Applications (IPA)
License: Creative Commons Attribution 4.0 International (CC BY 4.0)
Contact: Crestwood, IL, USA
This dataset is 100% synthetic and intended for educational use only.
It was generated from published physical models and real compactor specifications —
not measured on any real equipment, customer, or production batch. Do not use it for
regulatory… See the full description on the dataset page: https://huggingface.co/datasets/IPA-Marketing/Dry-Granulation_Roll-Compaction_and_Milling.wadar-soil-compactionfast-kv-compaction-cache
