Research Staff, Voice AI Foundations
Posted 1hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Research Staff pioneering latent-space models, neural audio codecs, and speech systems at Deepgram. Advancing scalable, efficient voice AI for real-time human-machine interaction.
Responsibilities:
- Pioneer Latent Space Models to address fundamental data, scale, and cost challenges in robust, contextualized voice AI
- Build next-generation neural audio codecs with extreme low-bit-rate compression and high-fidelity reconstruction
- Develop steerable generative models synthesizing diverse human speech, including emotional, multi-speaker, noisy, and overlapping-speech scenarios
- Create embedding systems that factorize latent spaces into speaker, content, style, environment, and channel dimensions
- Use latent recombination to generate synthetic audio data at previously impossible scales
- Train multimodal speech-to-speech systems for universal understanding and empathic, human-like responses
- Design model architectures, training schemes, and inference algorithms optimized for bare-metal hardware
- Enable cost-efficient training on billion-hour datasets and real-time inference for hundreds of millions of concurrent conversations
- Conduct rigorous controlled experiments, ablations, evaluations, and stress tests
- Collaborate through research publications and open-source contributions
Requirements:
- Strong mathematical foundation in statistical learning theory, particularly for self-supervised and multimodal learning
- Deep expertise in foundation model architectures and scaling training across multiple modalities
- Ability to derive novel mathematical formulations and implement them efficiently
- Demonstrated ability to build and maintain massive, high-quality, diverse data pipelines
- Experience designing controlled experiments to isolate architectural impacts and validate theoretical insights
- Experience optimizing models for real-world deployment, including hardware constraints and efficiency techniques
- Track record of open-source contributions or research publications advancing speech/language AI
- Ability to identify critical experiments that validate or disprove ideas quickly
- Vision to scale successful proofs-of-concept 100x
- Strong AI automation and augmentation mindset
- Scholarly literature or academic publications related to speech processing, STT/TTS, or similar requested in the application
- Legal authorization to work in the country where the role is located
- Ability to address visa sponsorship requirements for the country where the role is located
Benefits:
- Remote work arrangement
- AI-first work environment with opportunities to use and experiment with advanced AI tools
- Opportunity to work on transformative voice AI research
- Opportunity to contribute to open-source projects and research publications
- AI Notetaker interview recording is optional; opting out does not impact candidacy
















