Senior Site Reliability Engineer

Posted 1ds ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Senior SRE architecting multi-region cloud infrastructure, Kubernetes, disaster recovery, and observability for JumpCloud’s AI-powered IT management platform. Driving reliability, FinOps, automation, and incident management.

Responsibilities:

  • Architect, scale, and continuously improve the reliability, availability, and performance of JumpCloud’s multi-region microservices, APIs, and authentication infrastructure on AWS/GCP
  • Architect, build, and maintain Disaster Recovery processes, multi-region failover automation, and business continuity strategies
  • Lead the design and enforcement of SLIs, SLOs, and Error Budget frameworks
  • Drive end-to-end observability strategy using Datadog and Golden Signals monitoring
  • Lead on-call escalation and major incident management while driving adherence to 99.99% availability SLAs
  • Facilitate blameless post-incident reviews and implement systemic root-cause remediations
  • Architect, manage, and scale production Kubernetes EKS clusters with GitOps workflows using Argo CD and Kargo
  • Design and maintain Infrastructure-as-Code using Terraform across multi-account, multi-region cloud environments
  • Build FinOps and cost-optimization dashboards for multi-cloud spend, unit economics, and resource utilization
  • Write production-grade Python or Go tooling, platform automation, and custom integrations
  • Champion AI-assisted software development workflows using Cursor, Claude Code, and GitHub Copilot
  • Author operational runbooks and architecture decision records
  • Mentor mid-level and junior engineers

Requirements:

  • 8+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical, highly available distributed systems
  • Bachelor's degree in Computer Science, Software Engineering, or equivalent technical discipline
  • Advanced Python or Go capabilities
  • Hands-on experience with production EKS/GKE cluster lifecycles, ingress/egress, networking, RBAC, and GitOps tooling such as Argo CD
  • Deep Terraform proficiency across complex multi-account AWS environments, including IAM, VPCs, Transit Gateway, ALB/NLB, and Route53
  • Experience driving cloud cost-efficiency strategies, resource right-sizing, cost-allocation tagging, workload optimization, and FinOps dashboards
  • Experience designing and testing multi-region Disaster Recovery architectures and automating failover systems
  • Experience defining SLI/SLOs, managing PagerDuty schedules, and optimizing production observability platforms
  • Experience designing and operating enterprise service meshes such as Istio or Linkerd and production ingress/proxy systems such as HAProxy or NGINX
  • Ability to lead technical discussions, write architectural design documents/RFCs, and mentor engineering peers
  • Strong problem-solving, communication, and collaboration skills
  • Preferred: basic understanding of chaos engineering; secrets management architectures; DevSecOps practices; identity services, IAM, enterprise directory platforms, or security-focused SaaS solutions
  • Must be located in and authorized to work in India
  • Fluent spoken and written English required

Benefits:

  • Remote work within India
  • Opportunity to work with talent across 15+ countries
  • Professional growth and expertise-sharing opportunities
  • Supportive, collaborative work environment
  • Equal opportunity employment