Site Reliability Engineer

Posted 1ds ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Site Reliability Engineer strengthening AWS/GCP reliability, observability, Kubernetes, and disaster recovery. Supporting JumpCloud’s AI-powered unified IT management platform through automation and incident response.

Responsibilities:

  • Design, deploy, and maintain reliability, availability, and performance for critical JumpCloud systems and APIs across AWS and GCP
  • Operationalize SLIs, SLOs, and error budgets with core application teams
  • Build and refine end-to-end observability across microservices and cloud infrastructure using tools such as Datadog
  • Implement monitoring based on Golden Signals: latency, traffic, errors, and saturation
  • Participate in on-call rotations, incident response, and blameless post-incident reviews
  • Manage production Kubernetes EKS clusters using GitOps workflows such as Argo CD and Kargo
  • Provision and secure multi-cloud infrastructure using modular Terraform
  • Develop and maintain disaster recovery dashboards, runbooks, multi-region failover automation, and validation tests aligned with RTO/RPO targets
  • Write production-grade Python or Go scripts and automation tools to eliminate operational toil
  • Use AI-assisted development tools such as Cursor, Claude Code, and GitHub Copilot for scripting, runbook generation, and incident triage

Requirements:

  • 5+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical systems
  • Proficiency in Python or Go for SRE tools, custom automation, and cloud integrations
  • Production experience with Kubernetes, container orchestration, and GitOps pipelines such as Argo CD
  • Experience writing, maintaining, and modularizing Terraform configurations
  • Experience operating AWS workloads, including EKS, IAM, VPC networking, Route53, and ALB/NLB, or GCP workloads
  • Practical experience with FinOps, cost-allocation tagging, resource right-sizing, and cloud-spend dashboards
  • Experience building disaster recovery dashboards, running failover drills, and configuring monitoring for system health and recovery metrics
  • Practical experience with Datadog or similar, PagerDuty, alerting hygiene, and SLI/SLO frameworks
  • Operational experience configuring and troubleshooting production service meshes such as Istio and high-availability proxy solutions such as HAProxy or NGINX
  • Strong troubleshooting skills and track record of improving operational efficiency through code
  • Strong team-player orientation and alignment with company core values
  • Willingness and ability to participate in on-call shifts
  • Fluent spoken and written English
  • Preferred: experience with GitHub Actions or GitLab Pipelines
  • Preferred: basic understanding of chaos engineering or resilience testing
  • Preferred: familiarity with HashiCorp Vault, AWS Secrets Manager, or External Secrets Operator
  • Preferred: basic knowledge of DevSecOps tools and infrastructure-as-code vulnerability remediation

Benefits:

  • Remote-first work within India
  • Opportunity to work in a fast, SaaS-based environment
  • Professional growth and expertise-sharing opportunities
  • Collaboration with talented global teams
  • Employee voice in product and feature development
  • Supportive executive team and board
  • Equal opportunity employment
  • No third-party resumes accepted