COMMUNITY

BOARD

News

Regular Paper Accepted at the Top International Conference ASPLOS 2026

Author
College of Software Convergence
Date
2026-03-04
View
42

Files




A paper by master's student Song Geun-su (first author) and Professor Lee Young-min (corresponding author) has been accepted at the ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) 2026, the most prestigious international conference in systems software and architecture. ASPLOS is recognized by the Korean Institute of Information Scientists and Engineers as a top conference (BK recognized IF=4) and will be held in Pittsburgh, United States, from 22 to 26 March 2026.


The research, to be presented under the title 'oFFN: Outlier and Neuron-aware Structured FFN for Fast yet Accurate LLM Inference', finds that outliers appear concentrated at particular positions in the activation vector and, on that basis, proposes oFFN, a technique for accelerating LLM inference. oFFN statically rearranges weights in the feed-forward network (FFN) layer by taking account of both the activation frequency of output neurons and the frequency with which outliers appear. This exploits activation sparsity efficiently, achieving high inference acceleration with no loss of accuracy.


Accelerating inference by exploiting activation sparsity is a highly effective approach because it can relieve memory bottlenecks as well as computational ones, but accurately predicting which outputs will be sparse is a difficult problem. There is also the limitation that as batch size grows, structural sparsity falls and the acceleration benefit diminishes. By rearranging FFN weights to cluster outlier dimensions and to cluster neurons with similar sparsity efficiently, this research alleviates both problems at once. As a result it achieved acceleration of up to 5.46 times for the FFN and up to 2.01 times for total inference time (against a theoretical upper bound of 2.18 times) with almost no accuracy loss, a 13% improvement in inference speed over the previous state of the art.




Professor Lee Young-min said: "Exploiting activation sparsity is very promising for accelerating LLM inference, but output activation sparsity has been hard to exploit accurately and efficiently. oFFN overcomes that limitation by rearranging FFN weights on the basis of interesting observations about the characteristics of LLM inference and structurally clustering outlier dimensions and neurons. It is also significant as research that turns activation sparsity into real inference acceleration for multi-batch as well as single-batch inference, by using the GPU's tensor cores and CUDA cores complementarily. We plan to keep developing follow-up research in this area."

 

[References]

ASPLOS conference: https://www.asplos-conference.org/

High-Performance AI Systems Laboratory: https://aisys.sogang.ac.kr