Staff ML Platform Engineer

Posted 1ds ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Staff ML Platform Engineer building Datavant’s healthcare data and AI infrastructure. Leading production ML platforms, LLM serving, observability, and regulated-data safeguards.

Responsibilities:

  • Set technical direction across ML training, serving, and observability
  • Serve as final escalation point for complex infrastructure problems involving GPU capacity, Spark tuning, and production incidents
  • Own and evolve the paved-road framework, including shared CI/CD, model-workflow scaffolding, and Databricks Asset Bundles
  • Lead architecture for LLM endpoint serving across Databricks, AWS, Snowflake, and self-hosted deployments
  • Address latency, cost, caching, evaluation, and PHI-safe routing for LLM serving
  • Own standards and tooling for MLflow, model registry, training image supply chain, and observability
  • Partner with Data & ML Platform, Data Science, App Dev, and Operations teams
  • Provide technical input to vendor and platform selection decisions
  • Mentor senior engineers and guide platform consumers
  • Write high-leverage code and Infrastructure-as-Code hands-on

Requirements:

  • 10+ years of software engineering experience, including 3+ years designing, evolving, and operating enterprise-scale ML platforms in production
  • Strong technical judgment under ambiguity and track record of setting standards and influencing peers
  • Production experience with Databricks and/or Amazon SageMaker
  • Experience with MLflow or equivalent tracking and registry system
  • Experience with at least one core ML framework, such as PyTorch or TensorFlow
  • Fluency in Java or a JVM equivalent and Python
  • Deep Apache Spark experience for large-scale data and distributed compute
  • Deep AWS experience, including networking, IAM, GPU compute, storage, and messaging services
  • Fluency with Terraform, containers, Kubernetes, and GitHub-based CI/CD for ML workloads
  • Direct experience serving LLMs in production, including cost management, evaluation harnesses, and safe handling of sensitive prompts and outputs
  • Daily use of Claude Code, Cursor, Copilot, or equivalent AI coding tools
  • Clear written and verbal communication, especially in asynchronous remote settings
  • This job is not eligible for employment sponsorship
  • Preferred: technical leadership on healthcare or regulated-industry ML platforms; Databricks Asset Bundles, Unity Catalog, Iceberg, Delta; specialized inference pipelines; GPU capacity planning; Kafka or Kinesis; clinical or safety-sensitive AI evaluation and red teaming; open-source ML infrastructure or production ML publications

Benefits:

  • Total rewards compensation strategy
  • Post-offer health screenings and vaccinations as required by clients
  • Reasonable accommodations for individuals with physical and mental disabilities
  • Equal employment opportunity protections