Regular Paper Accepted at the Outstanding International Conference, The 30th International European Conference on Parallel and Distributed Computing (Euro-Par)
Files
The research team of the Machine Learning Systems Laboratory,
Regular Paper Accepted at the Outstanding International Conference, The 30th International European Conference on Parallel and
Distributed Computing (Euro-Par)

▶ (From left) Lee Eun-ji (master's), Han Yun-sang (master's), supervisor Professor Moon Eui-hyun
A paper written by graduate students of the Machine Learning Systems Laboratory — Lee Eun-ji of the Department of Computer Science & Engineering (master's, 3rd semester, joint first author) and Han Yun-sang of the interdisciplinary programme in Artificial Intelligence (master's, 3rd semester, joint first author) — with supervisor Professor Moon Eui-hyun (corresponding author) has been accepted at the international high-performance computing conference International European Conference on Parallel and Distributed Computing (Euro-Par). Euro-Par is listed at an adjusted IF of 1 among the outstanding international conferences in computer science under the BK21 Plus programme.
The paper, titled 'Accelerated Block-Sparsity-Aware Matrix Reordering for Leveraging Tensor Cores in Sparse Matrix-Multivector Multiplication', proposes a new algorithm for optimizing SpMM (Sparse Matrix-Multivector Multiplication) — a key kernel in many deep learning models and computational science applications — that maximizes the use of the Tensor Cores accelerators built into recent GPUs to speed up execution.

Figure 1. Overview of the parallelized sparse matrix reordering algorithm
To exploit Tensor Cores, designed to accelerate dense matrix–dense matrix operations, for SpMM acceleration, the paper proposes a sparse matrix reordering algorithm that converts a highly sparse matrix into a new data structure with high local density. Figure 1 shows the overall process of the parallelized sparse matrix reordering algorithm proposed in the paper. To cluster rows whose column-index patterns of non-zero elements are similar, the paper applies weighted Jaccard similarity to measure similarity between rows. This converts a sparse matrix with poor data locality into a new sparse matrix with high local density and improved data locality. Submatrices densely packed with non-zero elements in the newly converted sparse matrix are then designed to have their matrix multiplication accelerated through Tensor Cores. In addition, an optimized GPU kernel using dynamic parallelism was implemented so that the sparse matrix reordering preprocessing is accelerated even for huge sparse matrices with irregular sparsity patterns. Compared with the latest SpMM optimization algorithms and NVIDIA's sparse matrix operation library using Tensor Cores on the DLMC (Deep Learning Matrix Collection) dataset, which contains sparse matrices with a wide range of sparsity ratios and sizes, the BSA-SpMM (Block-Sparsity-Aware SpMM) algorithm proposed in the paper showed higher performance.
The 30th International European Conference on Parallel and Distributed Computing will be held in Madrid, Spain, from 26 to 30 August.