Staff ML Engineer – AWS Trainium, SageMaker

Posted 11hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Staff ML Engineer building production PyTorch training pipelines on AWS Trainium and SageMaker. Delivering cost-aware, hardware-optimized AI systems for enterprise clients.

Responsibilities:

  • Train and operate models on Amazon SageMaker with AWS Trainium as the underlying compute
  • Write and optimize PyTorch training code for Trainium, including reasoning about NeuronCore architecture, compiler behavior, memory, and throughput tradeoffs
  • Diagnose hardware-specific training run issues and distinguish data or code problems from compiler- or device-level problems
  • Translate Trainium job requests into working, cost-aware end-to-end training pipelines
  • Tune distributed training runs for throughput and cost on SageMaker training infrastructure
  • Work directly with client and internal engineering teams to scope and deliver production training workloads

Requirements:

  • Strong, hands-on PyTorch experience, ideally including distributed or multi-device training
  • Production experience with Amazon SageMaker for training and/or inference
  • Comfort working close to the hardware layer, including device-specific compilation and accelerator-level debugging
  • AWS Trainium or Inferentia (Neuron SDK) experience is a strong plus
  • Deep PyTorch experience and a track record of picking up new hardware targets quickly may substitute for Trainium experience
  • Solid Python fundamentals
  • Comfort operating in a client-facing, production engineering environment