SRE Engineer

Posted 1ds ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

SRE Engineer operando cloud, Kubernetes e observabilidade para a Verity, consultoria de transformação e engenharia digital. Prevenção de incidentes, automação e evolução de ambientes resilientes.

Responsibilities:

  • Define and track SLIs, SLOs, SLAs, MTTR, and MTTD
  • Implement observability, monitoring, alerting, and APM
  • Monitor latency, traffic, errors, saturation, availability, and performance
  • Work on incident prevention, identification, and resolution
  • Lead root cause analyses and define actions to prevent recurrence
  • Identify risks, bottlenecks, and single points of failure
  • Support the design of resilient, scalable, and highly available solutions
  • Automate operational activities and reduce manual tasks
  • Operate and evolve Kubernetes and Docker environments
  • Support capacity planning, business continuity, and disaster recovery strategies
  • Participate in deployments and support application stabilization
  • Collaborate with teams to improve reliability from the solution design stage
  • Create and maintain dashboards, alerts, procedures, and operational documentation
  • Promote a culture of reliability, observability, and continuous improvement

Requirements:

  • Experience as a Site Reliability Engineer, SRE, or in an equivalent role
  • Hands-on experience with cloud environments using GCP, AWS, and/or Azure
  • Knowledge of Kubernetes and Docker
  • Experience with observability, monitoring, alerting, and APM
  • Knowledge of SRE metrics and practices, such as SLI, SLO, SLA, MTTR, and MTTD
  • Experience managing, investigating, and resolving incidents
  • Knowledge of application and infrastructure troubleshooting
  • Experience administering Linux environments
  • Knowledge of networking, security, performance, and high availability
  • Experience with automation and Infrastructure as Code
  • Experience with CI/CD pipelines
  • Strong communication skills and the ability to work with cross-functional teams
  • Analytical, proactive, collaborative, and prevention-oriented mindset
  • Nice to have: experience with GKE, EKS, or AKS
  • Nice to have: knowledge of Dynatrace, Datadog, Grafana, Prometheus, or similar tools
  • Nice to have: experience with the ELK Stack, Elasticsearch, and Kibana
  • Nice to have: knowledge of Terraform and Ansible
  • Nice to have: experience with mission-critical environments and distributed systems
  • Nice to have: experience in financial institutions or regulated environments
  • Nice to have: experience with capacity management and cloud cost optimization
  • Nice to have: knowledge of disaster recovery and business continuity
  • Nice to have: experience defining and managing error budgets
  • Nice to have: certifications in Cloud, Kubernetes, or SRE

Benefits:

  • Meal voucher
  • Food allowance
  • Home office allowance
  • Health insurance
  • Dental insurance
  • Life insurance
  • Birthday Day Off
  • Total Pass / Wellhub app
  • Boon Saúde
  • Discount partnerships
  • Agreements with businesses and educational institutions
  • Welcome kit
  • Verity onboarding program
  • Verity Learning Interval
  • Great Place to Work certification and workplace improvement initiatives