Site Reliability Engineer
Posted 1ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Site Reliability Engineer strengthening AWS/GCP reliability, observability, Kubernetes, and disaster recovery. Supporting JumpCloud’s AI-powered unified IT management platform through automation and incident response.
Responsibilities:
- Design, deploy, and maintain reliability, availability, and performance for critical JumpCloud systems and APIs across AWS and GCP
- Operationalize SLIs, SLOs, and error budgets with core application teams
- Build and refine end-to-end observability across microservices and cloud infrastructure using tools such as Datadog
- Implement monitoring based on Golden Signals: latency, traffic, errors, and saturation
- Participate in on-call rotations, incident response, and blameless post-incident reviews
- Manage production Kubernetes EKS clusters using GitOps workflows such as Argo CD and Kargo
- Provision and secure multi-cloud infrastructure using modular Terraform
- Develop and maintain disaster recovery dashboards, runbooks, multi-region failover automation, and validation tests aligned with RTO/RPO targets
- Write production-grade Python or Go scripts and automation tools to eliminate operational toil
- Use AI-assisted development tools such as Cursor, Claude Code, and GitHub Copilot for scripting, runbook generation, and incident triage
Requirements:
- 5+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical systems
- Proficiency in Python or Go for SRE tools, custom automation, and cloud integrations
- Production experience with Kubernetes, container orchestration, and GitOps pipelines such as Argo CD
- Experience writing, maintaining, and modularizing Terraform configurations
- Experience operating AWS workloads, including EKS, IAM, VPC networking, Route53, and ALB/NLB, or GCP workloads
- Practical experience with FinOps, cost-allocation tagging, resource right-sizing, and cloud-spend dashboards
- Experience building disaster recovery dashboards, running failover drills, and configuring monitoring for system health and recovery metrics
- Practical experience with Datadog or similar, PagerDuty, alerting hygiene, and SLI/SLO frameworks
- Operational experience configuring and troubleshooting production service meshes such as Istio and high-availability proxy solutions such as HAProxy or NGINX
- Strong troubleshooting skills and track record of improving operational efficiency through code
- Strong team-player orientation and alignment with company core values
- Willingness and ability to participate in on-call shifts
- Fluent spoken and written English
- Preferred: experience with GitHub Actions or GitLab Pipelines
- Preferred: basic understanding of chaos engineering or resilience testing
- Preferred: familiarity with HashiCorp Vault, AWS Secrets Manager, or External Secrets Operator
- Preferred: basic knowledge of DevSecOps tools and infrastructure-as-code vulnerability remediation
Benefits:
- Remote-first work within India
- Opportunity to work in a fast, SaaS-based environment
- Professional growth and expertise-sharing opportunities
- Collaboration with talented global teams
- Employee voice in product and feature development
- Supportive executive team and board
- Equal opportunity employment
- No third-party resumes accepted


















