About Me¶
Personal Mission Statement¶
"Advancing human language technology and machine learning to deliver practical, impactful, and measurable improvements in everyday life."
Career Highlights¶
- Former Speech Architect of Voci’s speech recognition engine (2012-2023)
- Former maintainer of the open-source CMU Sphinx project (2004–2006)
- Early team member at two startups:
- ScanScout (#4 employee), acquired by Tremor Video in 2010
- Voci (#9 employee), later acquired by Medallia
- Research staff roles at Raytheon BBN, SpeechWorks (now part of Nuance), ScanScout, and Voci
- Co-author of a best paper at a prestigious international conference
Skills¶
Programming¶
- Languages: C, Perl, Python, C++, Java
- Also familiar with web and iOS programming
Conversational AI¶
My expertise spans the full stack of Conversational AI technologies, including automatic speech recognition (ASR), large language models (LLMs), and text-to-speech (TTS).
- ASR: Proficient across multiple generations of speech recognition systems, from traditional HMM-based models and HMM-DNN hybrids to modern foundation model–based approaches
- TTS: Experienced in fine-tuning state-of-the-art TTS foundation models for custom voice synthesis
- LLM: Deep experience with large language models, including prompt engineering, fine-tuning, and applying them to real-world NLP tasks.
🗣️ Speech Recognition¶
- System architecture (decoder + trainer)
- Very fast speech recognition
- Keyword spotting
- Classification tasks: topic, language, emotion, gender
- Robust speech recognition
Toolkits used:
PyTorch, kaldi, Sphinx (2, 3, 4, PocketSphinx), HTK, Julius, SpeechWorks (pre-OSR 2.0), Byblos, CMULM, SRILM, MITLM, KenLM
Deep Learning¶
- Large Language Model: Application of LLM on conversational systems, fine-tuning and RAG
- Speech Recognition: Low-level source-code experience with multiple deep learning toolkits
- Speech Synthesis: Fine-tuning CosyVoice2 for new language and accent correction.
- Administration: Setup and deployment of deep learning tools
Toolkits:
PyTorch, Tensorflow.
Soft Skills¶
- Startup experience and MVP building
- Speech applications and analytics
- Business use cases of speech recognition and machine learning (especially in startup contexts)
- Expertise in open-source speech recognition ecosystems