I specialize in developing Audio AI systems at scale.
Experienced in different areas like data processing, multilingual code-switched ASR, speaker diarization/verification, multimodal model, speech translation, source separation, TTS, and on-device models. Proven track record of translating research into production and open-source frameworks used across industry and academia. large language models, speaker diarization, and healthcare AI systems.
I am a Research Scientist at Zoom Communications working on production-scale Automatic Speech Recognition (ASR) for Voice AI Agents, healthcare speech systems, and word timing estimation.
Previously, I was a Postdoctoral Associate at the Massachusetts Institute of Technology (MIT CSAIL), where I worked with Dr. James Glass on robust speech recognition, trustworthy healthcare AI, uncertainty-driven pseudo-label filtering, and privacy-preserving voice deepfakes.
I completed my Ph.D. and M.S. in Computer Science at IIT Madras, working with Prof. C. Chandra Sekhar and Prof. Hema Murthy. My research focused on speaker diarization, continual transfer learning, and robust speech systems.
I have also contributed extensively to the open-source SpeechBrain toolkit as a core contributor during my internship at MILA, Montreal, Canada with Prof. Mirco Ravanelli and Prof. Yoshua Bengio.