Lead Site Reliability Engineer – Cloud

Posted 7hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Lead Site Reliability Engineer ensuring reliability and performance of Scalingo's cloud platform. Leading an SRE team and shaping practices for technical excellence in a growing tech startup.

Responsibilities:

  • Ensure the stability, availability and resilience of production systems
  • Anticipate failures and design effective incident response processes
  • Industrialize and automate platform operations
  • Maintain a high level of service quality for our customers and comply with contractual commitments (SLAs)
  • Analyze performance, identify bottlenecks and propose improvements to optimize resource usage and scalability
  • Define, implement and improve observability tools (monitoring, metrics, logs, alerting) with a proactive approach
  • Provide level-3 customer support in coordination with support teams and according to SLAs
  • Lead and facilitate incident retrospectives (post-mortems), identify root causes and define sustainable corrective actions

Requirements:

  • Strong expertise in cloud environments and distributed infrastructure
  • Proficiency in observability practices (logs, metrics, alerting) and a structured approach to diagnosing complex incidents
  • Solid understanding of containerized environments and their operational challenges
  • Proven skills with production databases: reliability, backups, restores, replication and scaling
  • Hands-on experience with Infrastructure as Code and environment automation
  • Awareness of operational security concerns
  • Comfortable using AI tools to improve day-to-day efficiency
  • Ability to operate in complex, changing or uncertain contexts with rigor and reliability
  • Comfortable prioritizing work, including during incident situations
  • Clear and structured communication, a preference for cross-functional collaboration and knowledge sharing
  • A blameless mindset, technical curiosity, composure and a focus on user impact
  • Ability to provide technical leadership, mentor others and advance collective practices

Benefits:

  • Fully remote with one trip per quarter (Strasbourg or another city)
  • Company events: one annual offsite and regular afterworks/social gatherings
  • Telework allowance (€57.60)
  • Meal vouchers (Ticket Restaurant) (€11.52 per voucher) and Swile card with its benefits
  • Flexible hours under a fixed-hours agreement (including RTT)
  • Linux laptop provided
  • Budget contribution for additional equipment