Senior Site Reliability Engineer, Production Engineering
Posted 4ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior Site Reliability Engineer supporting NVIDIA’s Cloud products through Kubernetes administration, automation, observability, and incident management. Maintaining highly available production services across global operations.
Responsibilities:
- Lead a global, dynamic, state-of-the-art Service Reliability Operations center
- Provide support for NVIDIA Cloud products and services
- Partner with Site Reliability Engineering, Security Operations Center, DevOps teams, and other organizations
- Support Production Kubernetes Services with a focus on automation and reducing manual tasks
- Perform large-scale Kubernetes administration, systems administration, and security monitoring to maintain service SLAs, integrity, and reliability
- Use alerts, alarms, and observability tools to monitor, detect, prevent, and respond to incidents
- Analyze logs, metrics, and system behavior to troubleshoot issues
- Lead root cause analysis and implement effective resolutions
- Initiate and lead incident management calls
- Coordinate subject matter experts and service owners for timely incident escalation and resolution
- Develop monitors, alarms, and alerts to improve service reliability and customer experience
Requirements:
- 7+ years of demonstrated experience administering large-scale production Kubernetes systems in high-availability Internet, Cloud, or Data Center environments
- Strong preference for on-prem expertise
- BS in Computer Science, Engineering, Mathematics, or equivalent experience
- Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management
- Familiarity with GPU / DPU hardware and high-performance computing Cluster environments
- Strong Linux system administration, DNS, DHCP and core Linux networking (IP Tables, routing, firewalls) experience
- Skills to troubleshoot and maintain services on large-scale bare-metal infrastructure
- Experience working with CI/CD tools like Jenkins, ArgoCD
- Experience in scripting
- Programming in Python or Golang or Rust preferred, but not required
- Strong communication and soft skills, able to present to cross-functional group members in a persuasive manner
Benefits:
- 24/7 Production engineering team support
- Flexibility to work on split-weekend shifts







