Staff ML Engineer – AWS Trainium, SageMaker
Posted 11hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Staff ML Engineer training and optimizing PyTorch models on AWS Trainium and SageMaker. Delivering cost-aware production AI pipelines for enterprise clients.
Responsibilities:
- Train and operate models on Amazon SageMaker with AWS Trainium as the underlying compute
- Write and optimize PyTorch training code, including understanding of NeuronCore architecture, compiler behavior, memory, and throughput tradeoffs
- Diagnose training run issues caused by hardware, distinguishing data or code problems from compiler- or device-level issues
- Translate Trainium job requests into working, cost-aware end-to-end training pipelines
- Tune distributed training runs for throughput and cost on SageMaker infrastructure
- Work directly with client and internal engineering teams to scope and deliver production training workloads
Requirements:
- Strong, hands-on PyTorch experience, ideally including distributed or multi-device training
- Production experience with Amazon SageMaker for training and/or inference
- Comfort working close to the hardware layer; understanding of device-specific compilation and ability to debug accelerator-related issues
- AWS Trainium or Inferentia (Neuron SDK) experience is a strong plus
- Deep PyTorch experience and a track record of picking up new hardware targets quickly if lacking Trainium/Inferentia experience
- Solid Python fundamentals
- Comfort operating in a client-facing, production engineering environment



















