Site Reliability Engineer
Posted 7hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Site Reliability Engineer scaling Arango’s cloud-native contextual AI data platform. Automating Kubernetes, AWS, and Google Cloud infrastructure with Golang, observability, and resilient production operations.
Responsibilities:
- Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms
- Ensure the scalability, performance, and reliability of Kubernetes-based distributed database systems
- Collaborate with developers to write production-grade Golang code for infrastructure automation and system operations
- Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems
- Develop disaster recovery, high availability, and fault-tolerance strategies
- Identify bottlenecks, troubleshoot, and resolve issues across networking, operating systems, and cloud infrastructure
- Implement monitoring, logging, and alerting systems
- Participate in on-call rotations and respond to production incidents
- Collaborate with cross-functional teams to improve reliability and scalability
- Collaborate with Customer Success to resolve customer issues
Requirements:
- Proven experience as an SRE or DevOps Engineer in a cloud-native environment
- At least 3 years with Kubernetes in a production environment
- Minimum 3 years of experience deploying and managing production-level cloud resources
- Proficiency with Kubernetes for large-scale distributed systems
- Experience with AWS and Google Cloud (GCP)
- Understanding of networking, security practices, and troubleshooting
- Understanding of Linux internals, including processes and environment variables
- Familiarity with Docker and containerization technologies
- Knowledge of CI/CD practices and tools such as Jenkins and CircleCI
- Familiarity with Prometheus, Grafana, ELK stack, and observability tools
- Knowledge of Git and version control systems
- Familiarity with Golang or Python, or willingness and capability to learn and apply Golang
- Ability to participate in on-call rotations
- Strong troubleshooting, problem-solving, communication, and collaboration skills
- Ability to self-organize and work independently as part of a remote team
- Nice-to-have: distributed database or large-scale data storage experience
- Nice-to-have: cloud security best practices
- Nice-to-have: Python or Bash scripting
- Nice-to-have: Terraform and Infrastructure-as-Code
- Nice-to-have: GitOps experience
- Nice-to-have: strong Golang programming experience
















