Regular Paper Accepted at the Outstanding International Conference International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS)
Files
Regular Paper Accepted at the Outstanding International Conference International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS)

Kim Ki-hyun, combined master's–doctoral student (supervisor: Professor Kim Young-jae)
The paper 'FLEXLLM: Flexible and Cost-Efficient LLM Serving with Heterogeneous GPUs', written by Kim Ki-hyun (combined master's–doctoral, first author) and Kim Jin-woo (combined master's–doctoral) of the Data-Centric Computing and AI Systems Laboratory (DISCOS) with Professor Dong Li (UC Merced) and Professor Kim Young-jae (corresponding author), has been accepted for publication at the outstanding international conference International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS).
Large language models (LLMs) have rapidly established themselves as a core technology across many fields, but the essential question that determines performance and efficiency in real inference environments is which architecture to choose. The KV cache that accumulates in large volumes during LLM inference rapidly consumes GPU memory, constraining batch size and context length, and simply adding more resources to relieve this offers limited value for money. Even with the same hardware configuration, performance varies greatly with the choice of architecture. Systematically selecting the optimal architecture for a given SLO (service level objective) and workload characteristics has therefore emerged as the central problem.

The figure above compares the structure and cost efficiency of three parallelization strategies (MP, AO, DP). On the left are the structures of model parallelism (MP), attention offloading (AO) and data parallelism (DP). MP distributes layers across multiple GPUs; AO performs the main computation on a primary GPU while offloading attention computation and KV cache storage to a secondary GPU; DP replicates the same model and splits the input batch. In the heatmap on the right, bold borders mark the optimal strategy for each GPU configuration, showing that even under identical conditions performance differs markedly with the strategy chosen. In particular, the KV cache offloading of the AO strategy reduces the burden on GPU memory, enabling long contexts or large batches.
FLEXLLM, developed by the team, integrates analysis of the user's SLO (service level objective), workload characteristics (input/output length, batch size) and the mix of heterogeneous GPUs to select automatically the optimal architecture among MP, AO and DP. The execution time of each strategy is estimated quickly with a lightweight prediction model and then linearly calibrated with a small number of measurements to align with real performance. The final configuration is then chosen from the candidates that satisfy the SLO on the basis of cost efficiency (tokens per dollar).
The experiments confirmed that on identical hardware, the choice of architecture alone improved cost efficiency (tokens per dollar) by up to 2.28 times. This clearly shows that precisely choosing the architecture that fits the problem and the objective, rather than adding expensive hardware, is what determines real service efficiency.
Kim Ki-hyun, the paper's first author, said: "I believe the direction of LLM inference is shifting from a speed race of 'bigger models and more expensive hardware' to how efficiently one can serve within a given SLO and cost constraint. Going forward I plan to extend beyond modeling into systems development that implements and improves real inference architectures integrating scheduling, batching and offloading. I am grateful to Kim Jin-woo, my fellow combined master's–doctoral student, for working on this with me and to Professor Kim Young-jae for his guidance, and I look forward to the active participation of interested undergraduates."
The International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication Systems (MASCOTS) is a distinguished conference on the modeling and performance analysis of computer systems. MASCOTS is listed at an adjusted IF of 2 among the outstanding international conferences in computer science under BK21, and as an outstanding conference in the Korean Institute of Information Scientists and Engineers' 2024 list of outstanding conferences in the software field. This year it will be held at Sorbonne University in Paris, France, from 21 to 23 October.
- The 33rd International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication Systems
- Website: https://mascots25.iitis.pl/