COMMUNITY

BOARD

News

Grand Prize in the Undergraduate Division at the Korea Computer Congress 2026 (KCC2026)

Author
College of Software Convergence
Date
2026-08-25
View
31

Files


In the undergraduate division of the Korea Computer Congress 2026 (KCC2026), held over three days from 24 June 2026, the paper 'Performance Analysis of Multi-GPU Parallelism in LLM Serving' by Choi Min-seok of the Department of Computer Science & Engineering (Computer Science & Engineering '21; supervisor: Professor Kim Young-jae) received the Grand Prize in the undergraduate division.

In large language model (LLM) serving, even with multiple GPUs, communication volume, KV cache headroom, batching efficiency and latency all vary with how the model and the requests are partitioned. This research compares tensor parallelism (TP), pipeline parallelism (PP), data parallelism (DP) and prefill-decode disaggregation (PD) under identical conditions in the same 2-GPU environment, analysing which parallelization scheme is advantageous and what structural bottlenecks arise according to input/output length and request load.

Synthetic workloads and the real BurstGPT trace were evaluated on a 2-GPU server with RTX 2080 Super 8GB cards using vLLM and Llama-3.2-1B-Instruct. In decode-heavy settings with long outputs, PD was superior in both throughput and ITL, and PD was also the best fit for batch-oriented workloads that prioritize throughput over response waiting time. In prefill-heavy settings with long inputs, however, the prefill bottleneck pushed average TTFT up to 4.71 seconds. In general settings with balanced input and output, TP offered the best balance of throughput and TTFT. TP gave the best balance of throughput and first-token latency on general workloads, while DP suffered KV cache pressure and request imbalance with long contexts and high concurrency, opening a GPU utilization gap of up to 94% and raising latency, making it unsuitable for services where latency stability matters. The work thus identifies the superior parallelization technique for each workload and demonstrates that the appropriate strategy must be chosen according to the service's priority metric.