ML Systems Engineer, WeLM

Chengxiang Qi 齐呈祥

I am a Machine Learning Systems Engineer on the WeLM team at Weixin Group (WXG), focusing on efficient LLM training and inference. My current work is on sparse attention for long-context training.

Previously at Microsoft Research Asia, I worked on FractalTensor and TileFusion. I am a co-first author of the FractalTensor paper at SOSP 2024.

I also build open-source systems tools and write about GPU programming and performance.
Chengxiang Qi

Selected work

TileFusion

☆ 119

Core designer & developer

A C++ template library for GPU tile computation. I helped design the abstractions and implement memory and compute primitives, enabling hardware-aware algorithms such as FlashAttention and FlashDecoding.

FractalTensor

☆ 32

SOSP 2024 · Co-first author

Exploring nested data parallelism and data reuse in deep learning. I implemented and optimized GPU computation with CUTLASS, with a reported average 2.14× speedup over the evaluated baselines on A100.

Light-DuoAttention

Long-context inference · CuTeDSL / SGLang

I implemented DuoAttention with CuTeDSL and integrated it into SGLang, connecting attention kernel development with a serving system.

Experience

July 2026 - Present

Weixin Group (WXG) / WeLM

Machine Learning System Engineer

Sparse attention for LLM training: Developing and optimizing sparse attention computation for long-context training of next-generation large language models, with a focus on kernel library development, numerical validation, and performance evaluation in preparation for integration into the training system.

June 2025 - June 2026

Weixin Group (WXG) / WeLM

Machine Learning System Intern

Built long-context inference kernels and integrated them into SGLang; explored GPU communication and distributed attention.

More about this role
  • Light-DuoAttention: Implemented Light-DuoAttention using CuTeDSL, achieving efficient long context inference and runs within SGLang.
  • NVSHMEM: Explored NVSHMEM and DeepEP and implemented the NVSHMEM-Tutorial, including hybrid communication based on CUDA IPC/RDMA, for technical sharing within the group.
  • Distributed Attention: Implemented Ring Attention Forward based on ThunderKittens using the LCF template, outperforming ring-flash-attention on short sequences. Implemented Flash Attention Backward based on LCF. Performed performance analysis on MagiAttention, ZigZag Ring Attention, and ZigZag Flex Attention.
  • Other: Investigated DSL on NVIDIA Hopper architecture.

Feb. 2024 - May 2025

Microsoft Research Asia

Research Intern in System and Network Group

Developed GPU implementations for FractalTensor and co-designed TileFusion, connecting programming abstractions with hardware-aware optimization.

More about this role
Mentor: Dr. Ying Cao
  • Based on the FractalTensor programming model, used CUTLASS to optimize the implementation of algorithms such as GEMM, Back-to-Back GEMMs, Stacked/Dilated LSTM, and FlashAttention-2. Performance evaluation was performed on NVIDIA A100, achieving up to 5.45x speedup compared with SOTA, with average performance acceleration of 2.14x.
  • As a core designer and developer, designed and implemented TileFusion, an efficient C++ macro kernel template library to improve the abstraction level of tile processing in CUDA C.

May 2023 - Aug. 2023

Tsinghua University

Research Intern in Operating System Lab

Developed Rust network drivers and virtualization support for Arceos, and worked on network performance.

More about this role
Mentors: Prof. Yu Chen, Dr. Yuekai Jia
  • Performed network performance benchmarking and optimization.
  • Wrote a driver for the Intel 82599 network interface card with Rust, referred to DPDK for performance optimization, and integrated it into Arceos.
  • Developed a type-2 hypervisor based on Arceos capable of booting Linux.

Publications

SOSP 2024 (*equal contribution)

Uncovering Nested Data Parallelism and Data Reuse in DNN Computation with FractalTensor

Siran Liu*, Chengxiang Qi*, Ying Cao, Chao Yang, Weifang Hu, Xuanhua Shi, Fan Yang, Mao Yang

Uncovering Nested Data Parallelism and Data Reuse in DNN Computation with FractalTensor
Earlier research · undergraduate thesis

Bachelor Thesis, Tianjin University (Outstanding Thesis Award)

基于 RISC-V 的 Type-1 Hypervisor 的设计与实现

Chengxiang Qi

基于 RISC-V 的 Type-1 Hypervisor 的设计与实现

Selected writing

All writings →

Open-source tools

ncu-cli ☆ 34

A Rust CLI tool for automated CUDA kernel performance diagnostics from NVIDIA Nsight Compute (NCU) CSV exports. It performs roofline analysis, architecture-aware heuristics, and profile diffing to generate actionable optimization suggestions in terminal, JSON, CSV, or Markdown.

Earlier projects & side projects 6

arceos-org/arceos ☆ 781

An experimental modular OS written in Rust. I integrated hypercraft into arceos, enabling it to be launched as a type-2 hypervisor. Added interrupt support and implemented IO interrupts based on virtio-net and virtio-blk. Implemented ixgbe NIC driver and performed network optimizations.

hypercraft ☆ 53

A VMM (Virtual Machine Monitor) crate written in Rust. Currently integrated into rcore-os/arceos and can be launched as a type-2 hypervisor capable of booting Linux. Developed from the hypocaust-2 project.

xv6-rust ☆ 347

Reimplementation of xv6-riscv in Rust language. A Unix-like operating system with optimizations to memory allocation and file system. Serves as the reference implementation for OSCOMP Project.

hypocaust-2 ☆ 60

A hardware-assisted virtualization RISC-V hypervisor using H extensions. Implements SBI call processing, two-stage page table translation, interrupt emulation and forwarding. Can boot & run rCore-Tutorial-v3, rt-thread and mainline Linux.

Demo

curgit ☆ 11

A high-performance CLI tool written in Rust that acts as a standalone Git Agent. It analyzes staged changes in a git repository and generates professional, context-aware commit messages following the Conventional Commits standard using LLM.

Where2Live ☆ 1

A Rust CLI tool to discover commutable residential communities near your workplace. Uses the Amap (Gaode) API to search for communities and rank them by transit commute time, rating, and popularity.

Education

Sep. 2023 - June 2026

University of Chinese Academy of Sciences

Master of Engineering in Computer Technology

Focused on deep learning compilers and machine learning systems.

Sep. 2019 - June 2023

Tianjin University

Bachelor of Engineering in Computer Science and Technology

Graduated with Outstanding Thesis Award.
Thesis: 基于 RISC-V 的 Type-1 Hypervisor 的设计与实现
Awards & talks
  • Outstanding Bachelor Thesis · Tianjin University · 2023
  • Third Prize in Team Competition · NSCSCC (National Student Computer System Capability Challenge) · 2022
  • The Best Quality Prize of OSPP · Open Source Promotion Plan (only 5 in China a year) · 2021
  • Third Prize in Team Competition · OSCOMP (Operating System Competition) · 2021

Hypocaust: a RISC-V type-1 hypervisor written in Rust

OS2ATC 2022, Beijing · March 2023

Presentation on the design and implementation of Hypocaust, a RISC-V type-1 hypervisor written in Rust, showcasing virtualization techniques and system-level Rust programming.

Slides

Beyond work

Away from the keyboard, I run and write. Here are my running log and personal essays.