TRACE: Tracking the Loss of Memory Provenance in LLM Agents

Year : 2026 | Volume : 13 | Issue : 02 | Page : 33 49
By

Jyoti Dabass,

Bhupender Singh Dabass,

  1. Global Professor of Practice,, Golden Gate University, California, USA
  2. Student, Department of Law, Institute of Law and Research, Haryana, India

Abstract

AI agents that carry memory across sessions gain better personalization and decision-making, but this persistence opens a serious security gap. When an agent repeatedly condenses earlier interactions into compact “lessons,” the trail linking each lesson back to the interaction that produced it gradually fades. We term this effect Reflective Attribution Collapse (RAC): the progressive loss of provenance and forensic traceability that results from repeated memory reflection. Under RAC, a malicious rule planted in memory can keep shaping agent behavior long after its origin becomes impossible to pinpoint or remove. Prior work has concentrated mainly on how attacks are injected and how they behave at retrieval time, leaving the forensic consequences of long-term summarization largely unexamined. We address this gap with TRACE (Targeted Rollback of Evolving Memory), a benchmark for quantifying attribution loss in reflective memory systems. TRACE relies on four measures: Lineage Recovery Rate (LRR), Provenance Decay Coefficient (PDC), Toxin Persistence Score (TPS), and Unattributable Attack Mass (UAM). In a clinical-agent environment built on MIMIC-III data, LRR falls from 91.3% to 14.7% over five reflection cycles while TPS stays above 70%, indicating that poisoned behavior can persist well after its source becomes effectively untraceable.

Keywords: Large language models (LLMs), AI agents, long-term memory, memory poisoning, reflective attribution collapse (RAC), provenance tracking, forensic traceability, memory security, retrieval-augmented systems, TRACE benchmark, AI safety, MIMIC-III

[This article belongs to Journal of Advancements in Robotics ]

How to cite this article: Jyoti Dabass, Bhupender Singh Dabass. TRACE: Tracking the Loss of Memory Provenance in LLM Agents. Journal of Advancements in Robotics. 2026; 13(02):33-49.
How to cite this URL: Jyoti Dabass, Bhupender Singh Dabass. TRACE: Tracking the Loss of Memory Provenance in LLM Agents. Journal of Advancements in Robotics. 2026; 13(02):33-49. Available from: https://journals.stmjournals.com/joarb/article=2026/view=259274

References

  1. Lin Z, Hao X, Fu R, Cui S, Chen K, Li C, et al. A survey on long-term memory security in LLM agents: attacks, defenses, and governance across the memory lifecycle [Preprint]. arXiv. 2026 Jun 11; arXiv:2604.16548v2 [cs.CR]. doi:10.48550/arXiv.2604.16548.
  2. Lam C, Li J, Zhang L, Zhao K. Governing evolving memory in LLM agents: risks, mechanisms, and the Stability and Safety Governed Memory (SSGM) framework [Preprint]. arXiv. 2026 May 19; arXiv:2603.11768v2 [cs.AI]. doi:10.48550/arXiv.2603.11768.
  3. Wen R, Li H, Xiao C, Zhang N. AgentSys: secure and dynamic LLM agents through explicit hierarchical memory management [Preprint]. arXiv. 2026 Feb 7; arXiv:2602.07398 [cs.CR]. doi:10.48550/arXiv.2602.07398.
  4. Bhardwaj VP. SuperLocalMemory: privacy-preserving multi-agent memory with Bayesian trust defense against memory poisoning [Preprint]. arXiv. 2026 Feb 17; arXiv:2603.02240 [cs.AI]. doi:10.48550/arXiv.2603.02240.
  5. Dong S, Xu S, He P, Li Y, Tang J, Liu T, et al. Memory injection attacks on LLM agents via query-only interaction. In: Belgrave D, Zhang C, Lin H, Pascanu R, Koniusz P, Ghassemi M, et al., editors. Advances in Neural Information Processing Systems. Vol. 38. Red Hook (NY): Curran Associates, Inc.; 2025. p. 46697-46731.
  6. Zou W, Dong M, Romero Calvo M, Chang S, Guo J, Lee D, et al. Poison once, exploit forever: environment-injected memory poisoning attacks on web agents [Preprint]. arXiv. 2026 Apr 7; arXiv:2604.02623v2 [cs.CR]. doi:10.48550/arXiv.2604.02623.
  7. Tian H, Sha Z, Wang J, Liu Y, Huang Z, Huang X. InjecMEM: memory injection attack on LLM agent memory systems [Preprint]. OpenReview. 2025 Sep 19. Available from: https://openreview.net/forum?id=QVX6hcJ2um
  8. Jing H, Li F, Dong Y, Zhou W, Liu R. Memory poisoning attacks on retrieval-augmented large language model agents via deceptive semantic reasoning. Eng Appl Artif Intell. 2026;167:113968. doi:10.1016/j.engappai.2026.113968.
  9. Dhivyasree T, Saravanan S, Ramamoorthy AK, Balasubramanian UM. Cognitive autonomous memory security (CAMS) against injection and extraction attacks in long-term memory of AI agents. Egypt Inform J. 2026;34:100983. doi:10.1016/j.eij.2026.100983.
  10. Yang S, Hu Z, Li X, Wang C, Yu T, Xu X, et al. DrunkAgent: stealthy memory corruption in LLM-powered recommender agents. In: Proceedings of the ACM Web Conference 2026 (WWW ’26). New York (NY): Association for Computing Machinery; 2026. p. 6853-6864. doi:10.1145/3774904.3792688.
  11. Deng X, Zhang Y, Wu J, Bai J, Yi S, Zou Z, et al. Taming OpenClaw: security analysis and mitigation of autonomous LLM agent threats [Preprint]. arXiv. 2026 Mar 12; arXiv:2603.11619 [cs.CR]. doi:10.48550/arXiv.2603.11619.
  12. Salem A, Paverd A, Abdelnabi S. Stateless yet not forgetful: implicit memory as a hidden channel in LLMs [Preprint]. arXiv. 2026 Feb 9; arXiv:2602.08563 [cs.LG]. doi:10.48550/arXiv.2602.08563.
  13. Bullwinkel B, Severi G, Hines K, Minnich A, Siva Kumar RS, Zunger Y. The trigger in the haystack: extracting and reconstructing LLM backdoor triggers [Preprint]. arXiv. 2026 Feb; arXiv:2602.03085 [cs.CR]. doi:10.48550/arXiv.2602.03085.
  14. Zhu H, Fiondella L, Yuan J, Zeng K, Jiao L. NeuroGenPoisoning: neuron-guided attacks on retrieval-augmented generation of LLM via genetic optimization of external knowledge. In: Belgrave D, Zhang C, Lin H, Pascanu R, Koniusz P, Ghassemi M, et al., editors. Advances in Neural Information Processing Systems. Vol. 38. Red Hook (NY): Curran Associates, Inc.; 2025. p. 73421-73446.
  15. Wu X, Ying L, Chen G, Gu Y, Qu H. Cache me, catch you: cache related security threats in LLM serving frameworks. In: Proceedings of the Network and Distributed System Security (NDSS) Symposium 2026; 2026 Feb 23-27; San Diego, CA, USA. San Diego (CA): Internet Society; 2026. doi:10.14722/ndss.2026.242812.
  16. Sunil BD, Sinha I, Maheshwari P, Todmal S, Mallik S, Mishra S. Memory poisoning attack and defense on memory based LLM-agents [Preprint]. arXiv. 2026 Jan 12; arXiv:2601.05504v2 [cs.CR]. doi:10.48550/arXiv.2601.05504.
  17. Torra V, Bras-Amorós M. Memory poisoning and secure multi-agent systems [Preprint]. arXiv. 2026 Mar 20; arXiv:2603.20357 [cs.CR]. doi:10.48550/arXiv.2603.20357.
  18. Jiang T, Wang Y, Liang J, Wang T. AgentLAB: benchmarking LLM agents against long-horizon attacks [Preprint]. arXiv. 2026 Feb 18; arXiv:2602.16901 [cs.AI]. doi:10.48550/arXiv.2602.16901.
  19. Muhammad S, Yusuf HM, Kumar S. Threats and defenses in large language models: a review of adversarial and model poisoning techniques [Preprint]. TechRxiv. 2026 Feb. doi:10.36227/techrxiv.177040550.02084667/v1.

Regular Issue Subscription Original Research
Volume 13
Issue 02
Received 19/05/2026
Accepted 14/06/2026
Published 27/06/2026
Publication Time 39 Days


Login

My IP

PlumX Metrics

Support