Regular Paper Accepted at the Outstanding International Conference, The IEEE International Symposium on Cluster, Cloud, and Internet Computing (CCGRID 2026)
Files

The paper 'LLM-Pilot: SLO-Aware and Cost-Efficient LLM Serving on Public Cloud VM Clusters via Offloading', written by Kim Jin-woo, combined master's–doctoral student at the Data-Centric Computing and AI Systems Laboratory (DISCOS) (first author; supervisor Professor Kim Young-jae), Kim Ki-hyun (combined master's–doctoral), Jung Hyun-seon (master's) and Professor Kim Young-jae (corresponding author), has been accepted for publication at The IEEE International Symposium on Cluster, Cloud, and Internet Computing (CCGrid 2026). A total of 247 papers were submitted this year, of which 62 were accepted as long regular papers (acceptance rate 25.1%).
As large language models (LLMs) such as ChatGPT, Claude and Gemini spread across industry, optimizing operating costs has emerged as a central challenge. In particular, as next-generation models such as Gemini 3 Pro and Claude Sonnet 4.5 support ultra-long contexts of more than a million tokens, GPU memory shortage has become a serious bottleneck. The KV cache generated during LLM inference grows linearly with sequence length, which translates directly into surging memory costs. The fundamental constraint of cloud environments, however, is structural: computing performance, memory capacity and I/O bandwidth are offered only as pre-packaged instances, so the resource actually needed cannot be scaled independently.

Figure 1. Overview of the LLM-Pilot methodology
LLM-PILOT, developed by the team, is a framework that systematically resolves this complex multi-dimensional optimization problem in the cloud. The system takes as input the workload characteristics (batch size, input/output sequence length), service level objectives (SLO: throughput, latency) and budget constraints, and optimizes the offloading strategy together with the combination of cloud instances. At the heart of the framework is a hierarchical analytical model. It abstracts LLM inference into three hardware-independent functions — 1. compute, 2. memory, 3. transfer — and mathematically predicts the performance bottleneck under each offloading strategy, such as KV cache offloading and attention offloading. A linear-regression-based calibration stage then mathematically reflects the differences between real hardware, further raising prediction accuracy.
Applied in a real AWS environment, LLM-Pilot achieves 2.05 times higher cost efficiency and 51% cost savings for real-time services compared with using a high-performance GPU alone, and 2.31 times higher cost efficiency for batch services compared with the latest prior work. In deterministic settings it achieves a mean absolute percentage error (MAPE) within 6% and a coefficient of determination (R2) above 0.89. Particularly notable is that in regions where existing methodologies are infeasible under strict SLO conditions, LLM-PILOT alone found a feasible configuration. This offers distinctive value beyond simple performance improvement, determining whether the service is possible at all.
Kim Jin-woo, the paper's first author and a combined master's–doctoral student, said the research greatly changed how he views LLM inference serving systems.
He said: "I came to feel that expensive equipment is not simply better; infrastructure that exactly fits the given requirements matters more. I realized that why you chose something is a more essential question than what you use."
He also said he was struck by the fact that, since the optimal cluster configuration changes with user requirements, budget and workload characteristics, LLM serving is not merely a systems implementation task but a design problem that must weigh economics and performance together. In cloud environments in particular, he added, such choices translate directly into cost, so he came to appreciate the importance of mathematical modeling and prediction-based decision-making.
IEEE CCGRID is a conference that aims to exchange fundamental advances and real-world applications in cloud computing, distributed systems, high-performance computing and artificial intelligence, to identify new research topics and to define the future of cloud computing. This year's event will be held in Sydney, Australia, from 18 to 21 May.
[References]
The 26th IEEE International Symposium on Cluster, Cloud, and Internet Computing (CCGRID`2026)
Website : https://ccgrid2026.cdms.westernsydney.edu.au/