Site Reliability Engineer

Posted 7hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Site Reliability Engineer scaling Arango’s cloud-native contextual AI data platform. Automating Kubernetes, AWS, and Google Cloud infrastructure with Golang, observability, and resilient production operations.

Responsibilities:

  • Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms
  • Ensure the scalability, performance, and reliability of Kubernetes-based distributed database systems
  • Collaborate with developers to write production-grade Golang code for infrastructure automation and system operations
  • Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems
  • Develop disaster recovery, high availability, and fault-tolerance strategies
  • Identify bottlenecks, troubleshoot, and resolve issues across networking, operating systems, and cloud infrastructure
  • Implement monitoring, logging, and alerting systems
  • Participate in on-call rotations and respond to production incidents
  • Collaborate with cross-functional teams to improve reliability and scalability
  • Collaborate with Customer Success to resolve customer issues

Requirements:

  • Proven experience as an SRE or DevOps Engineer in a cloud-native environment
  • At least 3 years with Kubernetes in a production environment
  • Minimum 3 years of experience deploying and managing production-level cloud resources
  • Proficiency with Kubernetes for large-scale distributed systems
  • Experience with AWS and Google Cloud (GCP)
  • Understanding of networking, security practices, and troubleshooting
  • Understanding of Linux internals, including processes and environment variables
  • Familiarity with Docker and containerization technologies
  • Knowledge of CI/CD practices and tools such as Jenkins and CircleCI
  • Familiarity with Prometheus, Grafana, ELK stack, and observability tools
  • Knowledge of Git and version control systems
  • Familiarity with Golang or Python, or willingness and capability to learn and apply Golang
  • Ability to participate in on-call rotations
  • Strong troubleshooting, problem-solving, communication, and collaboration skills
  • Ability to self-organize and work independently as part of a remote team
  • Nice-to-have: distributed database or large-scale data storage experience
  • Nice-to-have: cloud security best practices
  • Nice-to-have: Python or Bash scripting
  • Nice-to-have: Terraform and Infrastructure-as-Code
  • Nice-to-have: GitOps experience
  • Nice-to-have: strong Golang programming experience