Software Engineer III – Site Reliability
Posted 3ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Software Engineer III enhancing reliability and delivery systems at MyFitnessPal. Join a team focused on user experience and collaborative environments.
Responsibilities:
- Own and evolve our SLI/SLO and error-budget frameworks
- Lead incident response, drive postmortems
- Build and maintain observability across metrics, logs, and traces (Datadog)
- Design and operate resilient, scalable infrastructure using Infrastructure as Code (Terraform)
- Manage production Kubernetes and container workloads
- Own CI/CD pipelines and safe deployment strategies
- Own the security controls that live inside the delivery pipeline
- Implement and maintain policy-as-code
- Drive vulnerability triage and remediation SLAs for pipeline- and infrastructure-level findings
- Partner with our Security Engineer and the broader Security & Reliability disciplines
- Participate in and improve the on-call rotation
- Coach team members and engineers across the org on reliability patterns and operational best practices
Requirements:
- 5+ years in site reliability, platform, or infrastructure engineering, with clear senior-level ownership of production systems
- Strong programming skills for automation and tooling (Go, Python, Typescript or similar)
- Deep, hands-on experience with a major cloud platform (AWS is a plus), Kubernetes, and Infrastructure as Code (Terraform is a plus)
- Proven track record leading incident response and building SLO-driven reliability practices.
- Working fluency with observability tooling (Datadog is a plus)
- Practical experience integrating security into CI/CD pipelines — SAST/DAST/SCA tooling, dependency scanning, or policy-as-code
- Strong understanding of cloud security fundamentals
- The judgment and communication skills to raise a security or reliability finding with a senior engineer
- Experience with policy-as-code frameworks (especially Kyverno, but tools like OPA/Rego or Conftest are also relevant)
- Exposure to regulated or compliance-driven environments (SOC 2, PCI DSS, HIPAA) is a plus
- Chaos engineering or game-day experience is a plus
- Experience supporting B2C/mobile backend environments with high traffic, rapid iteration, and strong reliability needs is a plus
Benefits:
- healthcare
- parental planning
- mental health benefits
- annual performance bonus
- 401(k) plan and match
- responsible time off
- monthly wellness and technology allowances


















