Research Staff, Voice AI Foundations

Posted 1hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Research Staff pioneering latent-space models for codecs, speech generation, and multimodal voice AI. Advancing Deepgram’s real-time speech platform through scalable, efficient foundational research.

Responsibilities:

  • Develop next-generation neural audio codecs with extreme low-bit-rate compression and high-fidelity reconstruction across world-scale general audio corpora
  • Pioneer steerable generative models synthesizing diverse human speech, including emotional expression, multi-speaker scenarios, environmental noise, and overlapping speech
  • Develop embedding systems that factorize codec latent spaces into interpretable speaker, content, style, environment, and channel dimensions
  • Use latent recombination to generate synthetic audio data at unprecedented scales
  • Train multimodal speech-to-speech systems that understand diverse humans and produce empathic, human-like responses
  • Design model architectures, training schemes, and inference algorithms adapted to bare-metal hardware for cost-efficient billion-hour dataset training and real-time inference
  • Conduct foundational research in Latent Space Models addressing data, scale, and cost challenges in voice AI
  • Design controlled experiments, ablations, evaluations, stress tests, and benchmarks to validate research hypotheses
  • Collaborate through open-source contributions and research publications advancing speech and language AI

Requirements:

  • Strong mathematical foundation in statistical learning theory, particularly areas relevant to self-supervised and multimodal learning
  • Deep expertise in foundation model architectures and scaling training across multiple modalities
  • Proven ability to bridge theory and practice by deriving novel mathematical formulations and implementing them efficiently
  • Demonstrated ability to build data pipelines that process and curate massive datasets while maintaining quality and diversity
  • Track record of designing controlled experiments that isolate architectural innovations and validate theoretical insights
  • Experience optimizing models for real-world deployment, including hardware constraints and efficiency techniques
  • History of open-source contributions or research publications advancing speech/language AI
  • Ability to identify critical experiments that validate or disprove ideas quickly
  • Vision to scale successful proofs-of-concept 100x
  • Comfort using AI to automate and amplify personal impact
  • Ability to adapt quickly, experiment, learn constantly, and work in a rapidly changing AI environment

Benefits:

  • Remote work arrangement
  • AI-first work environment with active use and experimentation of advanced AI tools
  • Opportunity to pioneer foundational voice AI research with transformative impact
  • Opportunity to contribute to open-source projects and research publications
  • AI Notetaker interview recording is optional; opting out does not impact candidacy