I am currently a Ph.D. student at TCloud Lab, Shanghai Jiao Tong University (SJTU). I received my B.S. degree from Chien-Shiung Wu College, Southeast University.

Currently, my research is situated at the intersection of AI Infrastructure and High-Performance Computing. I am particularly interested in architecting efficient systems that scale deep learning workloads and optimizing performance for next-generation hardware. My goal is to bridge the gap between complex AI algorithms and the underlying hardware capabilities to enable more powerful and sustainable computing.

📝 Selected Publications

  • Eurosys 2025 Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores, Chenpeng Wu*, Qiqi Gu*, Heng Shi, et al.
    • Summary: This work addresses the compute and memory bottlenecks of MoE LLM inference by jointly exploiting structured sparsity in both activations and weights. It introduces a dedicated sparse data format and sparse-sparse kernels tailored for Sparse Tensor Cores, together with system-level optimizations for practical deployment.
  • PPoPP 2026 SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping, Qiqi Gu*, Chenpeng Wu*, Heng Shi, et al.
    • Summary: This paper targets redundant computation caused by zero-padding when mapping stencil workloads to Tensor Cores. SPIDER proposes a strided-swapping transformation (offline kernel transformation plus online input reordering) to satisfy 2:4 structured sparsity constraints and effectively enable Sparse Tensor Cores for stencil acceleration.
  • SC 2026 Do We Need Tensor Cores for Stencil Computations?, Qiqi Gu*, Chenpeng Wu*, Heng Shi, et al.
    • 🏆 Best Student Paper Nomination!
    • Summary: This study explains when and why Tensor Cores can accelerate stencil computations, which are traditionally viewed as memory-bound. It builds a performance model that accounts for transformation overheads and temporal fusion, identifying the conditions under which Tensor Core adaptation is beneficial.
  • TBD SptcAttn: Accelerating Long-Context Attention with Diagonal Stripes using Sparse Tensor Core, Chenpeng Wu*, Qiqi Gu*, Heng Shi, et al.
    • Summary: This work accelerates long-context sparse attention by mapping diagonal-stripe attention patterns to Sparse Tensor Core-compatible 2:4 sparsity. SptcAttn removes structured intra-tile redundancy without discarding selected attention scores, improving prefill performance while preserving model accuracy.
  • SC 2026 Pushing the Limits of Structured Sparse GEMM on Hopper GPUs via Analytical Modeling, Heng Shi, Qiqi Gu, Chenpeng Wu, et al.

* Equal contribution / Joint first authors.

📖 Educations

  • 2021.09 - (now), Pursuing Ph.D, Shanghai Jiao Tong University.
  • 2017.09 - 2021.06, Chien-Shiung Wu College, Southeast University.