SRE Engineer
Posted 1ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
SRE Engineer operando cloud, Kubernetes e observabilidade para a Verity, consultoria de transformação e engenharia digital. Prevenção de incidentes, automação e evolução de ambientes resilientes.
Responsibilities:
- Define and track SLIs, SLOs, SLAs, MTTR, and MTTD
- Implement observability, monitoring, alerting, and APM
- Monitor latency, traffic, errors, saturation, availability, and performance
- Work on incident prevention, identification, and resolution
- Lead root cause analyses and define actions to prevent recurrence
- Identify risks, bottlenecks, and single points of failure
- Support the design of resilient, scalable, and highly available solutions
- Automate operational activities and reduce manual tasks
- Operate and evolve Kubernetes and Docker environments
- Support capacity planning, business continuity, and disaster recovery strategies
- Participate in deployments and support application stabilization
- Collaborate with teams to improve reliability from the solution design stage
- Create and maintain dashboards, alerts, procedures, and operational documentation
- Promote a culture of reliability, observability, and continuous improvement
Requirements:
- Experience as a Site Reliability Engineer, SRE, or in an equivalent role
- Hands-on experience with cloud environments using GCP, AWS, and/or Azure
- Knowledge of Kubernetes and Docker
- Experience with observability, monitoring, alerting, and APM
- Knowledge of SRE metrics and practices, such as SLI, SLO, SLA, MTTR, and MTTD
- Experience managing, investigating, and resolving incidents
- Knowledge of application and infrastructure troubleshooting
- Experience administering Linux environments
- Knowledge of networking, security, performance, and high availability
- Experience with automation and Infrastructure as Code
- Experience with CI/CD pipelines
- Strong communication skills and the ability to work with cross-functional teams
- Analytical, proactive, collaborative, and prevention-oriented mindset
- Nice to have: experience with GKE, EKS, or AKS
- Nice to have: knowledge of Dynatrace, Datadog, Grafana, Prometheus, or similar tools
- Nice to have: experience with the ELK Stack, Elasticsearch, and Kibana
- Nice to have: knowledge of Terraform and Ansible
- Nice to have: experience with mission-critical environments and distributed systems
- Nice to have: experience in financial institutions or regulated environments
- Nice to have: experience with capacity management and cloud cost optimization
- Nice to have: knowledge of disaster recovery and business continuity
- Nice to have: experience defining and managing error budgets
- Nice to have: certifications in Cloud, Kubernetes, or SRE
Benefits:
- Meal voucher
- Food allowance
- Home office allowance
- Health insurance
- Dental insurance
- Life insurance
- Birthday Day Off
- Total Pass / Wellhub app
- Boon Saúde
- Discount partnerships
- Agreements with businesses and educational institutions
- Welcome kit
- Verity onboarding program
- Verity Learning Interval
- Great Place to Work certification and workplace improvement initiatives

















