Research Staff – Voice AI Foundations

Posted 1hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Research Staff developing latent-space models for Deepgram’s voice AI platform. Advancing neural audio codecs, generative speech, multimodal systems, and hardware-efficient inference.

Responsibilities:

  • Develop next-generation neural audio codecs with extreme low-bit-rate compression and high-fidelity reconstruction
  • Pioneer steerable generative models for diverse human speech, including emotional, multi-speaker, noisy, and overlapping-speech scenarios
  • Develop embedding systems that factorize latent representations into speaker, content, style, environment, and channel dimensions
  • Use latent recombination to generate synthetic audio data at previously impossible scales
  • Train multimodal speech-to-speech systems for robust understanding and empathic human-like responses
  • Design model architectures, training schemes, and inference algorithms optimized for bare-metal hardware
  • Enable cost-efficient training on billion-hour datasets and real-time inference for hundreds of millions of concurrent conversations
  • Conduct rigorous controlled experiments, ablations, evaluations, stress tests, and edge-case analyses
  • Contribute to foundational research advancing voice AI, speech, and language modeling

Requirements:

  • Strong mathematical foundation in statistical learning theory, particularly self-supervised and multimodal learning
  • Deep expertise in foundation model architectures and scaling training across multiple modalities
  • Ability to derive novel mathematical formulations and implement them efficiently
  • Demonstrated ability to build and curate massive datasets while maintaining quality and diversity
  • Track record of designing controlled experiments to validate architectural innovations and theoretical insights
  • Experience optimizing models for real-world deployment, including hardware constraints and efficiency techniques
  • History of open-source contributions or research publications advancing speech/language AI
  • Comfort actively using and experimenting with advanced AI tools
  • Ability to identify critical experiments that validate or disprove ideas quickly
  • Vision to scale successful proofs-of-concept 100x
  • Experience or knowledge relevant to neural audio codecs, generative models, embedding systems, latent representations, multimodal speech-to-speech systems, model architectures, training schemes, inference algorithms, and hardware-efficient deployment

Benefits:

  • Remote work arrangement
  • AI-first work environment with encouragement to use and experiment with advanced AI tools
  • Opportunity to work on transformative voice AI research at scale
  • Open-source contribution and research publication opportunities
  • AI Notetaker interview recording can be declined without affecting candidacy