Research Staff, Voice AI Foundations

Posted 1hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Research Staff pioneering latent-space models, neural audio codecs, and speech systems at Deepgram. Advancing scalable, efficient voice AI for real-time human-machine interaction.

Responsibilities:

  • Pioneer Latent Space Models to address fundamental data, scale, and cost challenges in robust, contextualized voice AI
  • Build next-generation neural audio codecs with extreme low-bit-rate compression and high-fidelity reconstruction
  • Develop steerable generative models synthesizing diverse human speech, including emotional, multi-speaker, noisy, and overlapping-speech scenarios
  • Create embedding systems that factorize latent spaces into speaker, content, style, environment, and channel dimensions
  • Use latent recombination to generate synthetic audio data at previously impossible scales
  • Train multimodal speech-to-speech systems for universal understanding and empathic, human-like responses
  • Design model architectures, training schemes, and inference algorithms optimized for bare-metal hardware
  • Enable cost-efficient training on billion-hour datasets and real-time inference for hundreds of millions of concurrent conversations
  • Conduct rigorous controlled experiments, ablations, evaluations, and stress tests
  • Collaborate through research publications and open-source contributions

Requirements:

  • Strong mathematical foundation in statistical learning theory, particularly for self-supervised and multimodal learning
  • Deep expertise in foundation model architectures and scaling training across multiple modalities
  • Ability to derive novel mathematical formulations and implement them efficiently
  • Demonstrated ability to build and maintain massive, high-quality, diverse data pipelines
  • Experience designing controlled experiments to isolate architectural impacts and validate theoretical insights
  • Experience optimizing models for real-world deployment, including hardware constraints and efficiency techniques
  • Track record of open-source contributions or research publications advancing speech/language AI
  • Ability to identify critical experiments that validate or disprove ideas quickly
  • Vision to scale successful proofs-of-concept 100x
  • Strong AI automation and augmentation mindset
  • Scholarly literature or academic publications related to speech processing, STT/TTS, or similar requested in the application
  • Legal authorization to work in the country where the role is located
  • Ability to address visa sponsorship requirements for the country where the role is located

Benefits:

  • Remote work arrangement
  • AI-first work environment with opportunities to use and experiment with advanced AI tools
  • Opportunity to work on transformative voice AI research
  • Opportunity to contribute to open-source projects and research publications
  • AI Notetaker interview recording is optional; opting out does not impact candidacy