Research Staff – Voice AI Foundations
Posted 1hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Research Staff developing latent-space models for Deepgram’s voice AI platform. Advancing neural audio codecs, generative speech, multimodal systems, and hardware-efficient inference.
Responsibilities:
- Develop next-generation neural audio codecs with extreme low-bit-rate compression and high-fidelity reconstruction
- Pioneer steerable generative models for diverse human speech, including emotional, multi-speaker, noisy, and overlapping-speech scenarios
- Develop embedding systems that factorize latent representations into speaker, content, style, environment, and channel dimensions
- Use latent recombination to generate synthetic audio data at previously impossible scales
- Train multimodal speech-to-speech systems for robust understanding and empathic human-like responses
- Design model architectures, training schemes, and inference algorithms optimized for bare-metal hardware
- Enable cost-efficient training on billion-hour datasets and real-time inference for hundreds of millions of concurrent conversations
- Conduct rigorous controlled experiments, ablations, evaluations, stress tests, and edge-case analyses
- Contribute to foundational research advancing voice AI, speech, and language modeling
Requirements:
- Strong mathematical foundation in statistical learning theory, particularly self-supervised and multimodal learning
- Deep expertise in foundation model architectures and scaling training across multiple modalities
- Ability to derive novel mathematical formulations and implement them efficiently
- Demonstrated ability to build and curate massive datasets while maintaining quality and diversity
- Track record of designing controlled experiments to validate architectural innovations and theoretical insights
- Experience optimizing models for real-world deployment, including hardware constraints and efficiency techniques
- History of open-source contributions or research publications advancing speech/language AI
- Comfort actively using and experimenting with advanced AI tools
- Ability to identify critical experiments that validate or disprove ideas quickly
- Vision to scale successful proofs-of-concept 100x
- Experience or knowledge relevant to neural audio codecs, generative models, embedding systems, latent representations, multimodal speech-to-speech systems, model architectures, training schemes, inference algorithms, and hardware-efficient deployment
Benefits:
- Remote work arrangement
- AI-first work environment with encouragement to use and experiment with advanced AI tools
- Opportunity to work on transformative voice AI research at scale
- Open-source contribution and research publication opportunities
- AI Notetaker interview recording can be declined without affecting candidacy
















