Encouragement Award in the Undergraduate Division at the Korea Software Congress 2025 (KSC2025)
Files

At the Korea Software Congress 2025 (KSC2025), held over three days from 16 December 2025, a paper by Jang Sung of the Department of Computer Science & Engineering (due to enter the master's course; supervisor: Kim Young-jae) received the Encouragement Award in the undergraduate division of the undergraduate/junior paper competition.
In LLM serving environments, differences in output length between requests cause head-of-line (HOL) blocking, a major cause of reduced overall throughput. The existing QLM predicts output length from past query statistics, but its prediction accuracy is low for out-of-distribution (OOD) queries that fall outside the training data distribution, causing HOL blocking and scheduling inefficiency.
To overcome this limitation, the research proposes S-QLM, which uses a small language model (SLM) to predict the output length of OOD queries in real time and reflects that in scheduling priority. In experiments on the ShareGPT dataset it achieved up to 11.5% higher throughput than the existing approach, demonstrating improved efficiency and stability in LLM serving environments.
