Regular Paper Accepted at the Top International Conference, The 38th International Conference for High Performance Computing, Networking, Storage, and Analysis (SC '26)
Files

The paper 'OASIS: Hierarchical Query Execution over Multi-Layer Computation-Enabled Object Storage', written by doctoral student Hwang Sun of the Data-Centric AI Computing and Systems Laboratory (DISCOS) (first author; supervisor Kim Young-jae), Professor Kim Young-jae (corresponding author), Park Jun-hyuk (master's), Yoo Jung-hyun (doctoral) and Ahn Seong-hoon (master's), together with the research team of SK hynix's Memory Systems Research (researchers Park Jung-an, Lee Jung-jin, Yang Jin-na and Yang Sun-yeol; team leader Noh Jung-ki; managers Jung Woo-seok and Kim Ho-sik) and researcher Qing Zheng of Los Alamos National Laboratory, has been accepted as a full paper at The 38th International Conference for High Performance Computing, Networking, Storage, and Analysis (SC '26). This year's acceptance rate for SC '26 full papers was 19.4%, so the work was accepted through fierce competition.
In large-scale scientific data analysis, moving tens of gigabytes of Parquet data from storage to the analysis cluster is a performance bottleneck. To alleviate this, computation-enabled object storage (COS) such as S3 Select and Skyhook has been proposed, reducing outbound data transfer by performing operators such as filters and projections on the storage server in advance. This approach, however, has the limitation that it does not reduce at all the data movement from the enclosure up to the storage server inside storage. In other words, the real bottleneck remains internal storage movement rather than external transfer. Existing systems stop at the binary decision of whether or not to execute an operator in storage, and do not address the essential question of where to place the execution boundary between multiple storage layers. Because resource constraints and data reduction ratios differ from layer to layer, placement in a single layer alone cannot optimize intermediate result size and inter-layer data movement.

Figure 1. Architecture overview and query processing flow of OASIS
This research formulates query execution as a hierarchical partitioning problem over a multi-layer stack of storage servers and enclosures. The key perspective is that the question is not which operators to offload, but where to cut and execute an already-offloaded query fragment — an execution boundary placement problem. This elevates storage-side query execution from a rule-based pushdown problem to a principled inter-layer optimization problem. The research starts from the conviction that a decomposition strategy must consider operator semantics, data reduction ratios and resource constraints together.
From this starting point the research proposes OASIS. OASIS is a unified execution framework spanning storage servers and enclosures; it expresses the query plan in a Substrait-based intermediate representation (IR), enabling fine-grained operator-level decomposition and flexible placement. At its core is SODA, a query decomposition algorithm that uses statistics built at ingestion time to estimate each operator's reduction ratio without reading data pages, determining the split point that minimizes inter-layer data movement. In the enclosure, SPDK and DuckDB are combined to perform file-system-free, column-aware execution, passing only the reduced intermediate results up to the higher layer. Maintaining an S3-compatible interface, so that compatibility with existing distributed SQL engines is preserved, is also central.
Evaluated on real scientific workloads including Laghos, DeepWater, CMS Open Data and TPC-H, OASIS reduced internal data movement from the enclosure to the storage server by up to 99.998% (20GB → 76KB for Q1). This lowered end-to-end latency by 11.4–25.1% compared with Skyhook, a state-of-the-art COS system, and achieved up to 47.6% higher performance on representative operator chains than a single-layer approach that executes only on the storage server. SODA's additional cost was negligible relative to end-to-end execution time: 0.126 ms for selectivity estimation and 1.81 ms for plan decomposition. The research also quantitatively established that offloading is not always beneficial and can be a loss in regions where there is no aggregation operator and output selectivity exceeds 25%.
Hwang Sun, the paper's first author and a doctoral student, said: "The usefulness of storage-side computation has long been raised, yet the resources inside the storage system itself have remained underused. By redefining this as a placement problem — at which layer inside storage should execution be cut — we set out to make the per-layer resources from the enclosure to the storage server actually used. As a result we were able to resolve the internal data movement bottleneck that had not previously been visible. I am grateful to Professor Kim Young-jae for his guidance and to the SK hynix and Los Alamos National Laboratory researchers who worked on this with us."
The International Conference for High Performance Computing, Networking, Storage, and Analysis (SC) is the world's most prestigious international conference in high-performance computing (HPC), where since its founding in 1988 researchers, system designers and industry experts have gathered to share the latest results in computing, storage and data analysis. It is a top conference that maintains a low acceptance rate through rigorous peer review each year; this year it will be held at McCormick Place in Chicago, United States, from 15 to 20 November, and accepted regular research papers will be presented in the Technical Program.
References:
- The 38th International Conference for High Performance Computing, Networking, Storage, and Analysis (SC '26)
- Website : https://sc26.supercomputing.org/