Staff ML Engineer – AWS Trainium, SageMaker
Posted 11hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Staff ML Engineer building production PyTorch training pipelines on AWS Trainium and SageMaker. Delivering cost-aware, hardware-optimized AI systems for enterprise clients.
Responsibilities:
- Train and operate models on Amazon SageMaker with AWS Trainium as the underlying compute
- Write and optimize PyTorch training code for Trainium, including reasoning about NeuronCore architecture, compiler behavior, memory, and throughput tradeoffs
- Diagnose hardware-specific training run issues and distinguish data or code problems from compiler- or device-level problems
- Translate Trainium job requests into working, cost-aware end-to-end training pipelines
- Tune distributed training runs for throughput and cost on SageMaker training infrastructure
- Work directly with client and internal engineering teams to scope and deliver production training workloads
Requirements:
- Strong, hands-on PyTorch experience, ideally including distributed or multi-device training
- Production experience with Amazon SageMaker for training and/or inference
- Comfort working close to the hardware layer, including device-specific compilation and accelerator-level debugging
- AWS Trainium or Inferentia (Neuron SDK) experience is a strong plus
- Deep PyTorch experience and a track record of picking up new hardware targets quickly may substitute for Trainium experience
- Solid Python fundamentals
- Comfort operating in a client-facing, production engineering environment



















