COMMUNITY

BOARD

News

Regular Paper Accepted at the Top International Conference ACM International Conference on Information and Knowledge Management (CIKM) 2026

Author
College of Software Convergence
Date
2026-08-27
View
40

Files



The paper 'AegisKV: Runtime Headroom Control for Online LLM Serving via Adaptive Attention Offloading', written by Kim Ki-hyun (combined master's–doctoral, first author), Jang Sung (combined master's–doctoral), Heo Min (master's), Kim Jin-woo (combined master's–doctoral) and Professor Kim Young-jae (corresponding author) of the Data-Centric Computing and AI Systems Laboratory (DISCOS), has been accepted for publication at the top international conference ACM International Conference on Information and Knowledge Management (CIKM) 2026 (acceptance rate: 597/2216 = 26.9%).

Large language models (LLMs) are now used in a variety of online services including conversational AI, coding assistance and AI agents. In settings that handle many requests at once or deal with long contexts, however, GPU memory and computing resources become a significant constraint. As the number of requests and the context length grow, the GPU memory needed during inference rises quickly, and when resources run short, running requests may be suspended or recomputation may occur to restore state. In the team's experiments, up to 43.1% of all GPU computation was spent on such recomputation in a memory-constrained GPU environment.

The usual remedies have been to add high-performance GPUs or to distribute model computation across several GPUs. These approaches are limited in build cost and resource efficiency, however, because expensive GPUs must be added even when only some resources are short.

With this in mind, the team developed AegisKV, a dynamic distribution technique that makes efficient use of heterogeneous GPUs with different performance and capacity within a single LLM inference system. AegisKV continuously tracks each GPU's computing performance, free memory and current processing state, and dynamically distributes the computation and data required for LLM inference according to each GPU's characteristics and the system load.

Unlike existing approaches that distribute every request uniformly across multiple GPUs, it analyses the resource usage of the running LLM workload in real time and selectively distributes only the work that needs distributing, when it needs it. This relieves the resource shortage on the primary GPU while minimizing the data movement and additional computation cost incurred by using a secondary GPU.

The core of AegisKV is not simply using several GPUs but dynamically adjusting the roles of GPUs with different performance and memory characteristics according to the state of the running LLM workload. To this end it monitors each GPU's resource state, dynamically decides what and how much to distribute, and efficiently overlaps inter-GPU data movement with computation to reduce additional latency. The team integrated these techniques into vLLM, a representative LLM serving framework, and implemented them as a working system.

In the experiments AegisKV delivered higher throughput than existing approaches across a variety of LLMs and online request settings. On large models in particular it shortened time to first token (TTFT) by up to 15.07 times, and adding a single secondary GPU raised total hardware cost by 4.8% while increasing the request volume that could be served by 42.9%.

The research shows that, rather than continually adding identical high-performance GPUs, selectively combining and dynamically using heterogeneous GPUs of differing performance and capacity according to the workload state can improve both the serving capacity and the cost efficiency of LLM services.

Kim Ki-hyun, the paper's first author and a combined master's–doctoral student, said: "This research is significant in approaching attention offloading not as a simple expansion of KV cache capacity but as the problem of dynamically using heterogeneous GPU resources according to the changing memory pressure of online LLM serving. We confirmed that rather than simply adding high-performance GPUs, using GPUs of different performance and capacity in roles that suit them can raise both service capacity and cost efficiency. I intend to extend this into research on LLM serving systems that use heterogeneous GPU resources efficiently by combining scheduling with memory management. I am grateful to my co-researchers Jang Sung, Heo Min and Kim Jin-woo, and to Professor Kim Young-jae for his guidance."

The ACM International Conference on Information and Knowledge Management (CIKM) is a globally prestigious international conference covering information retrieval, knowledge management, data mining, databases and artificial intelligence. CIKM is listed at a recognized IF of 3 among the outstanding international conferences in computer science under BK21, and is classified as a top conference in the Korean Institute of Information Scientists and Engineers' list of outstanding conferences in the software field. This year it will be held in Rome, Italy, from 7 to 11 November.

References:

●           35th ACM International Conference on Information and Knowledge Management (CIKM 2026)

●           Website: https://cikm2026.diag.uniroma1.it/