I am currently a Ph.D. student at TCloud Lab, Shanghai Jiao Tong University (SJTU). I received my B.S. degree from Chien-Shiung Wu College, Southeast University.
Currently, my research is situated at the intersection of AI Infrastructure and High-Performance Computing. I am particularly interested in architecting efficient systems that scale deep learning workloads and optimizing performance for next-generation hardware. My goal is to bridge the gap between complex AI algorithms and the underlying hardware capabilities to enable more powerful and sustainable computing.
📝 Selected Publications
Eurosys 2025Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores, Chenpeng Wu*, Qiqi Gu*, Heng Shi, et al.- Summary: This work addresses the compute and memory bottlenecks of MoE LLM inference by jointly exploiting structured sparsity in both activations and weights. It introduces a dedicated sparse data format and sparse-sparse kernels tailored for Sparse Tensor Cores, together with system-level optimizations for practical deployment.
PPoPP 2026SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping, Qiqi Gu*, Chenpeng Wu*, Heng Shi, et al.- Summary: This paper targets redundant computation caused by zero-padding when mapping stencil workloads to Tensor Cores. SPIDER proposes a strided-swapping transformation (offline kernel transformation plus online input reordering) to satisfy 2:4 structured sparsity constraints and effectively enable Sparse Tensor Cores for stencil acceleration.
SC 2026Do We Need Tensor Cores for Stencil Computations?, Qiqi Gu*, Chenpeng Wu*, Heng Shi, et al.- 🏆 Best Student Paper Nomination!
- Summary: This study explains when and why Tensor Cores can accelerate stencil computations, which are traditionally viewed as memory-bound. It builds a performance model that accounts for transformation overheads and temporal fusion, identifying the conditions under which Tensor Core adaptation is beneficial.
TBDSptcAttn: Accelerating Long-Context Attention with Diagonal Stripes using Sparse Tensor Core, Chenpeng Wu*, Qiqi Gu*, Heng Shi, et al.- Summary: This work accelerates long-context sparse attention by mapping diagonal-stripe attention patterns to Sparse Tensor Core-compatible 2:4 sparsity. SptcAttn removes structured intra-tile redundancy without discarding selected attention scores, improving prefill performance while preserving model accuracy.
SC 2026Pushing the Limits of Structured Sparse GEMM on Hopper GPUs via Analytical Modeling, Heng Shi, Qiqi Gu, Chenpeng Wu, et al.
* Equal contribution / Joint first authors.
📖 Educations
- 2021.09 - (now), Pursuing Ph.D, Shanghai Jiao Tong University.
- 2017.09 - 2021.06, Chien-Shiung Wu College, Southeast University.