Regular Paper Accepted at the Outstanding International Conference, The 28th Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD)
Files
The research team of the Machine Learning Systems Laboratory,
Regular Paper Accepted at the Outstanding International Conference, The 28th Pacific-Asia Conference on Knowledge Discovery and
Data Mining (PAKDD)
A paper written by graduate students of the Machine Learning Systems Laboratory — Yoon Bo-kyung of the Department of Computer Science & Engineering (master's, 4th semester, sole first author) and Han Yun-sang of the interdisciplinary programme in Artificial Intelligence (master's, 2nd semester) — with supervisor Professor Moon Eui-hyun (corresponding author) has been accepted at the international data mining conference Pacific-Asia Conference on Knowledge Discovery and Data Mining 2024 (PAKDD 2024). PAKDD is listed at an adjusted IF of 1 among the outstanding international conferences in computer science under the BK21 Plus programme.

(From left) Yoon Bo-kyung, Han Yun-sang, Professor Moon Eui-hyun
The paper, titled 'Layer-Wise Sparse Training of Transformer via Convolutional Flood Filling', proposes a new sparsification algorithm that reduces the computational complexity of multi-head attention — the bottleneck in transformer model computation — and thereby cuts the time and memory required for training.

Figure 1. Overview of the convolutional flood-filling algorithm in the sparse pattern generation/detection stage
Figure 1 shows the sparse pattern search algorithm proposed in the paper, which combines a convolution filter with a flood-filling algorithm to efficiently obtain the layer-wise sparse patterns that appear in attention computation. A convolution filter is first used to search for patterns in the attention score matrix. Average pooling then produces a block-level attention matrix, and the flood-filling algorithm identifies the important elements. Finally, upsampling generates a sparse pattern at the size of the attention score matrix. Training then resumes with sparse attention using the generated pattern. To reduce the training time and memory effectively, three operations of sparse attention were optimized. In experiments on six datasets, the method achieved better accuracy with up to 2.78 times less training time and 7.24 times less training memory than other sparsified transformer models.
The Pacific-Asia Conference on Knowledge Discovery and Data Mining 2024 will be held in Taipei, Taiwan, from 7 to 10 May.