Smaller Context, Recoverable History: Inside an Agent Memory Handoff
The Goal of a Memory Handoff
A memory handoff replaces a large conversation window while keeping retained source segments available for recall. The goal is to reduce how much history the LLM has to carry in its active context, while preserving the information it can recover when needed. That goal is different from shrinking a database. Keeping exact source text outside the prompt can be the right trade-off: the model stops processing the entire history on every turn, but retains a path back to the details. I checked the implementation, reran its targeted tests, and measured the replacement handoffs stored by the system. Across 10,241 completed runs, the selected source windows contained 1,137,387,720 characters. Their handoffs contained 15,375,395 characters in total-a 98.65% reduction in the replaced portion of context, measured in characters. I also recomputed the stored text hashes for 416,472 retained segments. None lacked exact source text, none had an exact or reduced hash mismatch, and none were missing a projection timestamp. I then ran a controlled synthetic comparison with a live answering model and real semantic embeddings. Full history and handoff-plus-recall both answered all 12 questions correctly. Handoff alone answered only the one question whose correct response was UNKNOWN. The recalled-context arm used substantially fewer input tokens, including the evidence retrieved for each answer. This is a small controlled result, not a claim of general lossless agent memory. The retrieval calls were made by the test harness, not chosen autonomously by an agent.
KEEP Does Not Mean "Keep in the Prompt"
KEEP does not mean "keep in the prompt". This is the most important detail in the implementation. The pipeline routes segments through three decisions:
- KEEP: Preserve exact text in external storage.
- COMPRESS: Store an accepted reduced representation alongside the exact source.
- DROP: Omit the segment from retained content.
All three decisions belong to the processing of a selected conversation window. Once its retained material has been durably handed off, that window can leave the active context-including its KEEP segments. The replacement is a compact handoff containing a short navigation summary, a manifest identifier, selected content hashes, and instructions for retrieving the omitted material. Consequently, a run with nothing but KEEP decisions can still substantially reduce active context. It retains information externally instead of repeatedly presenting all of it to the LLM.
What the Operational Measurements Show
What the operational measurements show. I queried aggregate lengths from completed runs without exporting user message text:
- Measurement: Result
- Completed runs: 10,241
- Selected source-window characters, summed: 1,137,387,720
- Replacement handoff characters, summed: 15,375,395
- Character reduction across those windows: 98.6482%
- Smallest handoff: 584 characters
- Largest handoff: 1,897 characters
The percentage is calculated from the summed lengths. It is not an average of per-run percentages. It also applies only to the replaced windows. System instructions, protected recent messages, tool definitions, and subsequent recall results still occupy context. The measurement is neither a whole-prompt reduction nor a token or inference-cost benchmark. Repeated runs can contain related material, so the source total is not a count of unique information. The earlier segment counters showed 416,459 KEEP decisions, 13 COMPRESS decisions, and no DROP decisions. The reduced storage representation was only 818 characters shorter than its source. That tiny difference answers a different question. It measures reduction within the externalized content, not the reduction from replacing the source window with its handoff. The large active-context change comes from externalization.
Comments
No comments yet. Start the discussion.