Research Focus
- High-performance kernels for low-bit, sparse, and hardware-aware computation, hand-tuned with Triton and CUDA for GPUs and FPGAs.
- Model compression through quantization, pruning, and knowledge distillation for large language models, MoE models, and generative models.
- Inference optimization, including KV cache and decoding strategies, expert routing/skipping, and attention/step caching for faster generation.
Ongoing Projects
- Clinical World Model for physiological forecasting.
- Real-time quantum error correction with customized FPGA.
- Deployment-friendly multimodal MoE compression.
- Super-resolution diffusion acceleration for image generation.
Selected Publications
All Publications
Filter Rules
Cleared!
Sort by Year