Profile Photo

Nauman Dawalatabad

Research Scientist at Zoom Communications, San Jose, CA, USA

I specialize in developing Audio AI systems at scale.

Experienced in different areas like data processing, multilingual code-switched ASR, speaker diarization/verification, multimodal model, speech translation, source separation, TTS, and on-device models. Proven track record of translating research into production and open-source frameworks used across industry and academia. large language models, speaker diarization, and healthcare AI systems.

About

I am a Research Scientist at Zoom Communications working on production-scale Automatic Speech Recognition (ASR) for Voice AI Agents, healthcare speech systems, and word timing estimation.

Previously, I was a Postdoctoral Associate at the Massachusetts Institute of Technology (MIT CSAIL), where I worked with Dr. James Glass on robust speech recognition, trustworthy healthcare AI, uncertainty-driven pseudo-label filtering, and privacy-preserving voice deepfakes.

I completed my Ph.D. and M.S. in Computer Science at IIT Madras, working with Prof. C. Chandra Sekhar and Prof. Hema Murthy. My research focused on speaker diarization, continual transfer learning, and robust speech systems.

I have also contributed extensively to the open-source SpeechBrain toolkit as a core contributor during my internship at MILA, Montreal, Canada with Prof. Mirco Ravanelli and Prof. Yoshua Bengio.

Experience

Research Scientist
Zoom Communications Inc.
2024 – Present
  • Production-grade ASR systems for Voice AI Agents in French and Hindi-English code-mixed settings.
  • Zoom's first Healthcare-focused ASR systems for doctor-patient conversations using LLM persona based synthetic TTS data and added contextual biasing. Used for generating clinical notes.
  • Lexicon-free multilingual word boundary estimation model used as core module for diarization, ASR, and EOU detection.
Postdoctoral Associate
Massachusetts Institute of Technology - Computer Science & Artificial Intelligence Laboratory
2021 – 2024
  • Worked with Dr. James Glass on robust ASR and healthcare AI systems.
  • Uncertainty-driven pseudo-label data filtering and calibration for speech recognition.
  • Multimodal dementia detection from long-form interviews.
  • Privacy-preserving voice deepfake generation for healthcare data augmentation.
  • Selected as IEEE ICASSP Rising Stars in Signal Processing (2023).
Lead Engineer — Bixby Speech & Audio Team
Samsung
2021
Research Intern — SpeechBrain Core Team
Mila / University of Montreal
2020 – 2021

Selected Publications

Differential Privacy Preserving Voice Conversion for Audio Health Data
Nauman Dawalatabad, Meysam Ahangaran, Cody Karjadi, Lampros Kourtis, Philip Joung, Vijaya B. Kolachalama, Rhoda Au, James Glass
Alzheimer's & Dementia, 2025
Cross-Lingual Transfer Learning for Low-Resource Speech Translation
Sameer Khurana, Nauman Dawalatabad, Antoine Laurent, Luis Vicente, Pablo Gimeno, Victoria Mingote, James Glass
IEEE ICASSP, 2024
On Unsupervised Uncertainty-Driven Speech Pseudo-Label Filtering and Model Calibration
Nauman Dawalatabad, Sameer Khurana, Antoine Laurent, James Glass
IEEE ICASSP, 2023
Detecting Dementia from Long Neuropsychological Interviews
Nauman Dawalatabad, Yuan Gong, Sameer Khurana, Rhoda Au, James Glass
EMNLP Findings, 2022
ECAPA-TDNN Embeddings for Speaker Diarization
Nauman Dawalatabad, Mirco Ravanelli, François Grondin, Jenthe Thienpondt, Brecht Desplanques, Hwidong Na
INTERSPEECH, 2021
Novel Architectures for Unsupervised Information Bottleneck based Speaker Diarization of Meetings
Nauman Dawalatabad, Srikanth Madikeri, C. Chandra Sekhar, Hema A. Murthy
IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), vol. 29, 2021

Skills

Speech Recognition
Speaker Diarization
Large Language Models
SpeechBrain
PyTorch
ESPNet
Kaldi
Healthcare AI
Voice Privacy
Python
TensorFlow

Some Invited Talks & Recognitions

Featured in News