Senior Site Reliability Engineer, AWS Cloud
Posted 2ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior Site Reliability Engineer architecting reliable, scalable AWS and multi-cloud platforms for Nagarro, a digital product engineering company. Automating operations, improving observability, and leading incident response and reliability engineering.
Responsibilities:
- Architect and drive the reliability, scalability, and performance of the multi-cloud provisioning platform across all production stacks
- Architect and implement end-to-end automation pipelines to eliminate manual intervention and reduce technical toil
- Define, monitor, and improve critical system health indicators (SLIs/SLOs), including latency, throughput, error rates, and capacity usage
- Make data-driven architectural recommendations
- Lead collaboration with product and cross-functional engineering teams to embed reliability and security considerations early into the SDLC
- Design robust incident detection mechanisms and manage incident response
- Triage critical incidents and lead deep Root Cause Analysis (RCA)
- Implement long-term preventative engineering solutions
- Audit platform operations to identify bottlenecks, eliminate single points of failure, and reduce structural complexity
Requirements:
- 8+ years of experience working within cloud environments in SRE or Cloud Platform/Reliability Engineer roles
- Strong experience in cloud development and multi-cloud environments
- Strong exposure to AWS cloud
- Knowledge of cloud architecture, scalability, and high-availability design
- Hands-on experience with Kubernetes and container orchestration
- Experience with Terraform and Infrastructure as Code (IaC)
- Experience designing automation and CI/CD pipelines to reduce operational toil
- Strong understanding of SRE principles, SLIs/SLOs, monitoring, and observability
- Proven experience with Incident Management, Root Cause Analysis (RCA), and reliability engineering
- Ability to identify performance bottlenecks, single points of failure, and architectural risks

















