Home / Current Issue / Paper 1718160
M.A.K.S: Multidimensional Access Knowledge Scoring for Long-Horizon LLM Agent Memory Management
Subject area: Science,Engineering and Technology · Area of research: Machine Learning
DOI: https://doi.org/10.64388/IREV9I11-1718160
Abstract
AI agents with long horizons suffer from fixed context window sizes that necessitate memory evictions over time. Current techniques such as FIFO, LRU, and attention-based evictions use a binary approach to manage memory by either retaining or irreversibly deleting data. No current system preserves its evicted memories for later recovery, nor do any of the systems use multiple criteria to determine the value of memories. M.A.K.S., which stands for Multidimensional Access Knowledge Scoring, is a memory management technique designed specifically for LLM agent systems. M.A.K.S uses a continuous memory lifecycle consisting of degradation and revivals to address the issue. Each memory has an associated priority score denoted by S(t) that takes into account several factors including temporal degradation, access frequency, centrality, Shannon Entropy, and spaced reinforcement. To assess M.A.K.S, we conducted three experiments. Through the experiment called the Needle in a Compressed Haystack, we show that M.A.K.S was able to successfully reconsolidate a very important memory fact with a token usage ratio of 4.58× where FIFO memory evictions failed completely. The ablation experiment corroborates the inclusion of the Ghost Zone and reconsolidation pipeline in architectural support structures. Overhead benchmarking proves that lazy evaluation improves sweep performance up to 9.25× faster with 5,000 memory units, ensuring that scores remain less than 1% of LLM inference latency. M.A.K.S shows that memory management of AI agents is a first-class systems issue involving scoreable degradation, cold storage, and cue-based revival – not simple eviction.
References
[1] Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., Wang, Z., and Chen, B. (2023). H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models. arXiv:2306.14048.
[2] Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M. (2023). Efficient Streaming Language Models with Attention Sinks. arXiv:2309.17453.
[3] Munkhdalai, T., Faruqui, M., and Gopal, S. (2024). Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention. arXiv:2404.07143.
[4] Packer, C., Fang, V., Patil, S. G., Moon, K., Wooders, S., and Gonzalez, J. E. (2023). MemGPT: Towards LLMs as Operating Systems. arXiv:2310.08560.
[5] Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023). Generative Agents: Interactive Simulacra of Human Behavior. arXiv:2304.03442.
[6] Ebbinghaus, H. (1885). Über das Gedächtnis: Untersuchungen zur experimentellen Psychologie. Duncker & Humblot, Leipzig. [Memory: A Contribution to Experimental Psychology. Translated by Ruger, H. A. and Bussenius, C. E., Teachers College, Columbia University, 1913.]
[7] Averell, L. and Heathcote, A. (2011). The form of the forgetting curve and the fate of memories. Journal of Mathematical Psychology, 55(1), 25–35.
[8] Leitner, S. (1972). So lernt man lernen: Der Weg zum Erfolg [How to Learn to Learn: The Way to Success]. Herder, Freiburg im Breisgau.
[9] Wozniak, P. A. and Gorzelanczyk, E. J. (1994). Optimization of repetition scheduling with the algorithm SM-2. Acta Neurobiologiae Experimentalis, 54(6), 59–67.
[10] Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., and Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380.
[11] Nader, K., Schafe, G. E., and Le Doux, J. E. (2000). Fear memories require protein synthesis in the amygdala for reconsolidation after retrieval. Nature, 406(6797), 722–726.
[12] Schacter, D. L., Guerin, S. A., and St. Jacques, P. L. (2012). Memory distortion: an adaptive perspective. Trends in Cognitive Sciences, 15(10), 467–474. [Reviewed in computational contexts as a model of trace reactivation and reconsolidation.]
[13] Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., Dong, Y., Tang, J., and Li, J. (2023). LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding. arXiv:2308.14508.
[14] Shaham, U., Segal, E., Levy, I., Yu, X., Groeneveld, J., Zettlemoyer, L., and Srikumar, V. (2022). SCROLLS: Standardized CompaRison Over Long Language Sequences. arXiv:2201.03533.
[15] Shannon, C. E. (1948). A Mathematical Theory of Communication. Bell System Technical Journal, 27(3), 379–423.
How to cite this paper
@article{1718160,
author = {Sahil Mehraj, Abdul Kafeel, Sheikh Musa},
title = {M.A.K.S: Multidimensional Access Knowledge Scoring for Long-Horizon LLM Agent Memory Management},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {11},
pages = {3864-3883},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1718160.pdf},
abstract = {AI agents with long horizons suffer from fixed context window sizes that necessitate memory evictions over time. Current techniques such as FIFO, LRU, and attention-based evictions use a binary approach to manage memory by either retaining or irreversibly deleting data. No current system preserves its evicted memories for later recovery, nor do any of the systems use multiple criteria to determine the value of memories. M.A.K.S., which stands for Multidimensional Access Knowledge Scoring, is a memory management technique designed specifically for LLM agent systems. M.A.K.S uses a continuous memory lifecycle consisting of degradation and revivals to address the issue. Each memory has an associated priority score denoted by S(t) that takes into account several factors including temporal degradation, access frequency, centrality, Shannon Entropy, and spaced reinforcement. To assess M.A.K.S, we conducted three experiments. Through the experiment called the Needle in a Compressed Haystack, we show that M.A.K.S was able to successfully reconsolidate a very important memory fact with a token usage ratio of 4.58× where FIFO memory evictions failed completely. The ablation experiment corroborates the inclusion of the Ghost Zone and reconsolidation pipeline in architectural support structures. Overhead benchmarking proves that lazy evaluation improves sweep performance up to 9.25× faster with 5,000 memory units, ensuring that scores remain less than 1% of LLM inference latency. M.A.K.S shows that memory management of AI agents is a first-class systems issue involving scoreable degradation, cold storage, and cue-based revival – not simple eviction.},
month = {May},
doi = {https://doi.org/10.64388/IREV9I11-1718160}
}