COMMUNITY

BOARD

News

Regular Paper Accepted at the Outstanding International Conference, The 40th International Conference on Massive Storage Systems and Technology (MSST 2026)

Author
College of Software Convergence
Date
2026-05-28
View
31

Files



The paper 'Scout-and-Reduce: Budget-Aware Storage-Side Reduction for Memory-Safe GPU Analytics', written by doctoral student Hwang Sun (first author) of the Data-Centric AI Computing and Systems Laboratory (DISCOS) and Professor Kim Young-jae (corresponding author) together with the research team of SK hynix's Memory Systems Research (team leader Noh Jung-ki, manager Jung Woo-seok), has been accepted as a full paper at The 40th International Conference on Massive Storage Systems and Technology (MSST 2026).

GPUs can process interactive tabular data analysis quickly through GPU dataframe libraries such as cuDF, and are therefore used as the core accelerator in workflows that analyse Parquet data held in object storage directly from a notebook environment. This direct path, however, exhibits a GPU memory cliff: as soon as the data grows a little, GPU memory suddenly blows up and the job fails outright. A job that completes normally at one input size fails at a larger scale with an out-of-memory (OOM) error and little warning. Because Parquet uses column encoding and page compression, memory usage after decoding is far larger than the size on disk, so OOM is hard to predict from file size alone.


Figure 1. Architecture overview of Scout-and-Reduce (S&R)


The research focuses on two budgets that determine the memory safety of GPU analytics. The first is Btransfer, the decoded size of the transfer unit brought onto the GPU in one step; the second is Bmerge, the size of the intermediate hash table state the GPU must maintain during group-by aggregation. Existing storage-side offloading techniques focus solely on reducing the number of bytes transferred and are limited in that they properly control neither budget. The research starts from the conviction that object storage should not merely deliver bytes but should emit units already shaped to fit the GPU's memory capacity.

From this starting point the research proposes Scout-and-Reduce (S&R). S&R consists of a sidecar co-located on the storage server and a GPU client; when the client sends its GPU memory information along with the query, the sidecar's two-stage planner runs. Reading only the footer metadata of the Parquet file, the planner estimates post-decoding GPU memory usage without reading any data pages at all, and selects the direct path if the job fits within GPU capacity or the S&R path, which performs exact local reduction on the storage side, if it does not. For high-cardinality group-by in particular, it partitions the results into key-disjoint buckets so that the GPU merges only one bucket at a time, guaranteeing that the intermediate state does not exceed device capacity. The key point is that it redefines the storage–client interface while preserving the semantics of the data itself.

Evaluated on a 4GB GPU (NVIDIA RTX A400) with two real datasets — NYC Yellow Taxi and OTTO recommendation events — S&R completed the job at every scale where the direct path and filter offloading failed with OOM. The footer-based estimator predicted the actual decoded size within an average error of 1.8%, and S&R kept peak GPU memory comfortably below the device limit regardless of data scale. All of this was performed using only the CPU and memory resources of the existing storage server, with no dedicated hardware, demonstrating its general applicability.

Hwang Sun, the paper's first author and a doctoral student, said: "Analytics jobs suddenly failing because of GPU memory limits is something one encounters often, yet it has not been treated in depth. I think the perspective of shaping data on the storage side to fit the GPU's memory budget can be a new direction for solving this problem. I am grateful to Professor Kim Young-jae for his guidance and to the SK hynix researchers who worked on this with us."

The International Conference on Massive Storage Systems and Technology (MSST) is a prestigious international conference where, since its founding in 1974, designers, architects, researchers and industry experts in large-scale storage systems have gathered to share the latest research results and challenges in file and storage systems. This year it will be held at Santa Clara University in Silicon Valley, United States, from 1 to 5 June, and peer-reviewed regular research papers will be presented in the research track (4–5 June).

References:

- The 40th International Conference on Massive Storage Systems and Technology (MSST 2026)

- Website : https://www.msstconference.org/