Research

PhD Candidate at Beihang University, specializing in neural network compression and model quantization. Making deep learning more efficient and accessible for real-world deployment.

Research Focus

  • High-performance kernels for low-bit, sparse, and hardware-aware computation, hand-tuned with Triton and CUDA for GPUs and FPGAs.
  • Model compression through quantization, pruning, and knowledge distillation for large language models, MoE models, and generative models.
  • Inference optimization, including KV cache and decoding strategies, expert routing/skipping, and attention/step caching for faster generation.

Ongoing Projects

  • Clinical World Model for physiological forecasting.
  • Real-time quantum error correction with customized FPGA.
  • Deployment-friendly multimodal MoE compression.
  • Super-resolution diffusion acceleration for image generation.

Selected Publications

All Publications

Filter Rules Cleared!
Sort by Year
View publications on Google Scholar →