Reach out: arkadipmaitra at gmail dot com
Working on distributed AI systems at Red Hat, contributing upstream to the PyTorch ecosystem. The work spans multi-node communication, training correctness, on-device inference, and research in computer vision and sign language.
Distributed training and communication form the core focus — fixing correctness and reliability issues in collective operations, transport layers, and multi-backend coordination for large-scale GPU clusters. This includes work on failure-path determinism, lifecycle management, and making communication behave predictably across hardware topologies.
On-device and edge inference work involves extending operator coverage, fixing kernel correctness, and improving build-system plumbing for selective compilation on resource-constrained targets.
Inference diagnostics and tooling contributions center on building source-level tracing and adapter infrastructure that enables any framework built on PyTorch to expose its internal call paths for debugging and profiling.
Developer tooling and framework integration work includes compiler interactions with distributed subsystems, autograd correctness for tensor subclasses, and improving error messages to surface real problems faster.
Active in technical reviewing for journals and certification programs, conference program committee work, mentoring at hackathons, giving talks at meetups and universities on distributed systems and PyTorch internals, and published technical writing on distributed communication infrastructure.