Regular Paper Accepted at the Outstanding International Conference, 34th International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS '26)
Files

The paper 'PAVE: Partial KV Cache Eviction for Fine-Grained KV Reclamation in LLM Serving', written by master's student Jang Sung of the Data-Centric AI Computing and Systems Laboratory (DISCOS) (first author; supervisor Kim Young-jae), Professor Kim Young-jae (corresponding author), Kim Ki-hyun (combined master's–doctoral), Kim Jin-woo (combined master's–doctoral) and Heo Min (master's), together with Professor Dong Li of the University of California, Merced, has been accepted as a full paper at the 34th International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS '26). This year's acceptance rate for MASCOTS '26 full papers was 23%, so the work was accepted through fierce competition.
As generative AI services such as ChatGPT spread rapidly, LLM serving technology that handles many users' requests quickly and simultaneously is growing in importance. An LLM stores the conversation the user has entered so far in GPU memory as a KV cache (key-value cache), reducing repeated computation when generating subsequent tokens. As conversations grow longer, however, the KV cache keeps growing, and when many users use the service at once GPU memory runs short quickly. To prevent this, existing LLM serving systems temporarily suspend a running request and remove that request's KV cache from GPU memory entirely (preemption).
The existing approach, however, has the limitation that it removes a request's entire KV cache even when memory is only slightly short. Because everything is deleted when only a small amount of memory needs to be freed, the work already computed must be recomputed from scratch when the request is served again. This unnecessary recomputation increases response latency and degrades the service level objective (SLO). Existing research has focused on when to preempt, which request to choose, or how quickly to restore the deleted KV cache, but has not addressed the fundamental question of how much KV cache should be removed.

Figure 1. Design overview of PAVE
This research defines the problem as over-eviction — excessive removal of the KV cache. That is, it notes that existing LLM serving systems remove far more KV cache than is actually needed, causing unnecessary recomputation. To resolve this, the research proposes PAVE (Partial KV Cache Eviction). Instead of removing a request's entire KV cache, PAVE is a new memory management technique that selectively reclaims only the KV blocks needed to free the currently insufficient memory. When the request runs again, only the deleted portion needs to be restored, greatly reducing recomputation. PAVE also has the advantage of being easy to apply to existing systems, since it improves only the memory reclamation process without changing the scheduler or computation structure of existing LLM serving systems.
Evaluated across a range of settings including the real conversational workload Azure Conversation and synthetic workloads, PAVE completely eliminated the excessive KV cache removal that occurred with the existing approach and reduced the average latency of segments requiring recomputation from more than 300 ms to about 113 ms. It also improved time per output token (TPOT), which reflects perceived performance, raising the SLO satisfaction rate by up to 1.87 times over vLLM, a representative LLM serving framework. These results show that PAVE can improve the performance of a variety of LLM serving systems while operating independently of the existing scheduling policy.
Jang Sung, the paper's first author and a master's student, said: "Being able to solve something existing research had not resolved and contribute to the field in my first semester of graduate school gave me a real sense of achievement and showed me the appeal of research. There were difficulties along the way, but I am sincerely grateful to Professor Kim Young-jae, who guided me to the end in every way, and to my seniors Kim Ki-hyun (combined master's–doctoral), Kim Jin-woo (combined master's–doctoral) and Heo Min (master's) for their unstinting help."
The IEEE International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS) is a long-established international conference on the performance analysis of computer systems and networks. Now in its 34th edition, it will be held at the University of Genoa in Genoa, Italy, from 20 to 22 October.
References:
- 34th International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication System (MASCOTS '26)
- Website : https://mascots26.iitis.pl/