Primary focus
Distributed training & collective communication
I work on the behavior that determines whether multi-node training is trustworthy: collective operations, transport layers, process-group lifecycle, backend coordination, and failure handling across heterogeneous hardware topologies.
The standard is simple: failures should be bounded and explainable, and successful work should preserve the correctness guarantees users expect.