Senior Site Reliability Engineer
Posted 1ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior SRE architecting multi-region cloud infrastructure, Kubernetes, disaster recovery, and observability for JumpCloud’s AI-powered IT management platform. Driving reliability, FinOps, automation, and incident management.
Responsibilities:
- Architect, scale, and continuously improve the reliability, availability, and performance of JumpCloud’s multi-region microservices, APIs, and authentication infrastructure on AWS/GCP
- Architect, build, and maintain Disaster Recovery processes, multi-region failover automation, and business continuity strategies
- Lead the design and enforcement of SLIs, SLOs, and Error Budget frameworks
- Drive end-to-end observability strategy using Datadog and Golden Signals monitoring
- Lead on-call escalation and major incident management while driving adherence to 99.99% availability SLAs
- Facilitate blameless post-incident reviews and implement systemic root-cause remediations
- Architect, manage, and scale production Kubernetes EKS clusters with GitOps workflows using Argo CD and Kargo
- Design and maintain Infrastructure-as-Code using Terraform across multi-account, multi-region cloud environments
- Build FinOps and cost-optimization dashboards for multi-cloud spend, unit economics, and resource utilization
- Write production-grade Python or Go tooling, platform automation, and custom integrations
- Champion AI-assisted software development workflows using Cursor, Claude Code, and GitHub Copilot
- Author operational runbooks and architecture decision records
- Mentor mid-level and junior engineers
Requirements:
- 8+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical, highly available distributed systems
- Bachelor's degree in Computer Science, Software Engineering, or equivalent technical discipline
- Advanced Python or Go capabilities
- Hands-on experience with production EKS/GKE cluster lifecycles, ingress/egress, networking, RBAC, and GitOps tooling such as Argo CD
- Deep Terraform proficiency across complex multi-account AWS environments, including IAM, VPCs, Transit Gateway, ALB/NLB, and Route53
- Experience driving cloud cost-efficiency strategies, resource right-sizing, cost-allocation tagging, workload optimization, and FinOps dashboards
- Experience designing and testing multi-region Disaster Recovery architectures and automating failover systems
- Experience defining SLI/SLOs, managing PagerDuty schedules, and optimizing production observability platforms
- Experience designing and operating enterprise service meshes such as Istio or Linkerd and production ingress/proxy systems such as HAProxy or NGINX
- Ability to lead technical discussions, write architectural design documents/RFCs, and mentor engineering peers
- Strong problem-solving, communication, and collaboration skills
- Preferred: basic understanding of chaos engineering; secrets management architectures; DevSecOps practices; identity services, IAM, enterprise directory platforms, or security-focused SaaS solutions
- Must be located in and authorized to work in India
- Fluent spoken and written English required
Benefits:
- Remote work within India
- Opportunity to work with talent across 15+ countries
- Professional growth and expertise-sharing opportunities
- Supportive, collaborative work environment
- Equal opportunity employment


















