Site Reliability Engineer

Posted 10hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Site Reliability Engineer maintaining Kubernetes and cloud infrastructure for Arango’s contextual AI data platform. Automating operations, observability, CI/CD, and reliability for enterprise AI systems.

Responsibilities:

  • Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms
  • Ensure scalability, performance, and reliability of Kubernetes-based distributed database systems
  • Collaborate with developers to write production-grade Golang code for infrastructure automation and system operations
  • Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems
  • Develop disaster recovery, high availability, and fault-tolerance strategies
  • Identify system bottlenecks and troubleshoot issues across networking, operating systems, and cloud infrastructure
  • Implement monitoring, logging, and alerting systems for system health and performance visibility
  • Participate in on-call rotations and respond to production incidents
  • Collaborate with cross-functional teams to improve system reliability and scalability
  • Collaborate with the Customer Success team to resolve customer issues

Requirements:

  • Proven experience as an SRE or DevOps Engineer in a cloud-native environment
  • Minimum 3 years of Kubernetes experience in a production environment
  • Minimum 3 years of experience deploying and managing production-level cloud resources
  • Proficiency with Kubernetes for large-scale distributed systems
  • Experience with AWS and Google Cloud (GCP)
  • Understanding of networking, security practices, troubleshooting methods, and Linux internals
  • Familiarity with Docker and containerization technologies
  • Knowledge of CI/CD practices and tools such as Jenkins and CircleCI
  • Familiarity with monitoring, alerting, and observability tools such as Prometheus, Grafana, and ELK stack
  • Knowledge of version control systems, particularly Git
  • Familiarity with Golang or Python; willingness and capability to learn and apply Golang accepted
  • Ability to self-organize and work independently as part of a remote team
  • Experience managing distributed databases or large-scale data storage systems is a plus
  • Knowledge of cloud security best practices is a plus
  • Experience with Python or Bash scripting is a plus
  • Terraform/IaC experience is a plus
  • GitOps experience is a plus
  • Strong Golang programming skills and automation-tool development experience are a plus

Benefits:

  • Remote work
  • Collaboration with experienced engineers, marketers, and product leaders
  • Opportunity to contribute to cutting-edge AI and data infrastructure
  • Opportunity to help shape how enterprises build AI-powered applications
  • Inclusive, growth-oriented team environment