Narmeen07/agent-session-handoff
Agent Session Handoff Synthetic operations-agent transcripts for studying knowledge retention across a model swap: does a compressed KV-cache memory keep more of a session than a text summary when the next model takes over? Each episode is built from a sampled fact record (service owners, ports, branches, config values, ticket states, decisions) rendered by an LLM writer into a realistic user / assistant / tool-output session. The writer never sees the questions. A swap point… See the full description on the dataset page: https://huggingface.co/datasets/Narmeen07/agent-session-handoff.
Agent Session Handoff
Synthetic operations-agent transcripts for studying knowledge retention across a model swap: does a compressed KV-cache memory keep more of a session than a text summary when the next model takes over?
Each episode is built from a sampled fact record (service owners, ports, branches, config values, ticket states, decisions) rendered by an LLM writer into a realistic user / assistant / tool-output session. The writer never sees the questions. A swap point separates the history (what gets compressed) from updates (new facts plus one ownership correction that the receiving model reads as plain text). Five questions per episode are generated from the record and scored by exact match:
Two configurations: default (pool v1: person names from 12 first × 12 last names per split, split-disjoint) and pool_v2 (person names composed from syllables, thousands per split, split-disjoint by hash; ids end in -p2). v1 was found to let a compressor memorise the name vocabulary; v2 is the recommended training configuration and the v1 test split doubles as an unseen-vocabulary test.
Controls built in: entity names are fictional and drawn from split-disjoint pools (no parametric leakage); n_facts (8 / 16 / 32) and filler_turns (8 / 24 / 48) vary independently so fact density and length can be swept; distractor facts are true but never asked; every subject and value is verified to appear verbatim, and the abstention subject is verified absent.
Fields
id,split,n_facts,n_distractors,filler_turns,seed,renderer,writer_modelhistory— the session transcript before the swapupdates— the post-swap text (three team memberships, one new port, one ownership correction)questions— list of{kind, question, answer, metric}record— the ground-truth facts, distractors and updates as[relation, subject, object]triples
Provenance
Generated by the cache-transfer-knowledge-injection project's schema-grounded session generator (Martian), writer model glm-4.5-air, seed 17. Split-disjoint name pools; train/dev come from different seeds streams than test by construction.
train: 2924 episodesdev: 297 episodestest: 300 episodes
This is a synthetic diagnostic set. Success on it does not establish real-world agent performance.
