COMMUNITY

BOARD

News

Regular Paper Accepted at the Outstanding International Conference IEEE International Symposium on Cluster, Cloud, and Internet Computing (CCGRID) 2026

Author
College of Software Convergence
Date
2026-02-25
View
34

Files



The paper 'QUICKTOPIA: Iteration-Level GPU Frequency Control for Energy-Latency Co-Optimization in LLM Inference', written by Baek So-yang (first author, master's graduate), Jung Bo-don (joint first author, doctoral), Dr Byun Hong-su and Professor Park Sung-yong (corresponding author) of the Data-Centric AI Computing and Systems Laboratory (DISCOS), has been accepted for publication at The IEEE International Symposium on Cluster, Cloud, and Internet Computing (CCGRID) 2026. A total of 247 papers were submitted this year, of which 62 were accepted as regular papers (acceptance rate 25.1%).

 

Deployment environments for large language models (LLMs) are rapidly expanding beyond large data centres to resource-constrained settings such as single-GPU servers, on-premises infrastructure and edge devices, and the power consumption and operating costs incurred during inference have accordingly emerged as a central bottleneck. Research using GPU DVFS (Dynamic Voltage and Frequency Scaling) has been pursued to alleviate this, but existing techniques mostly adjust frequency only within the range that satisfies explicit service level objectives (SLO), or rely on static profiling. Real LLM inference engines, however, see computation fluctuate sharply from iteration to iteration because of continuous batching; existing techniques do not sufficiently reflect these dynamics and, by focusing on either latency or energy, fail to achieve comprehensive optimization in terms of the energy-delay product (EDP).



Figure 1. The tree structure of QUICKTOPIA and its Pareto-optimality-based EDP optimization process


To overcome these limitations, the research proposes QUICKTOPIA, a GPU frequency control framework that analyses the iteration-level execution characteristics of LLM inference in real time and dynamically searches for the optimal balance between energy and latency. QUICKTOPIA adapts immediately to runtime variability without modifying the model architecture or performing offline profiling, and evaluates the optimal trade-off between the conflicting objectives of latency and energy in real time on the basis of Pareto optimality, a multi-objective optimization concept. It also designs a tree-based segment manager that minimizes control overhead and maintains a stable frequency policy even in noisy environments, maximizing practical system efficiency.

 

Integrating QUICKTOPIA into vLLM, a state-of-the-art LLM inference framework, and evaluating it in a single-GPU environment held latency within 1–5% across a range of workloads while reducing EDP, the comprehensive efficiency metric, by up to 22% compared with the latest prior techniques. This empirically demonstrates that QUICKTOPIA is a practical and effective LLM inference optimization solution that secures both energy efficiency and responsiveness across LLM services.

 

Baek So-yang, the paper's first author and a master's graduate, said: "I believe that as AI models advance, the systems software that supports them will become even more important. I hope this research on balancing performance and energy will have a positive influence on the many undergraduates and junior researchers interested in AI and systems software."

 

IEEE CCGRID is a world-class international conference in cluster, cloud and internet computing, and a prestigious venue for sharing the latest results in large-scale distributed systems, cloud infrastructure, high-performance computing and data-intensive systems research. This year's event will be held in Sydney, Australia, from 18 to 21 May.