Principal Site Reliability Engineer

Posted 3hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Principal SRE ensuring reliable cloud infrastructure for Tandem Diabetes Care’s insulin technology. Leading incident response, automation, disaster recovery, and compliance across distributed teams.

Responsibilities:

  • Lead day-to-day production support, including intake, triage, prioritization, escalation, queue health, and change execution
  • Establish consistent support practices across a distributed team, including shift handoffs, ticket quality standards, and ownership of open issues
  • Lead incident management end-to-end, including incident command, stakeholder communication, and blameless postmortems with corrective actions tracked to closure
  • Coordinate production security incident response with Security teams
  • Own on-call strategy, rotation design, escalation paths, alert tuning, and tooling such as PagerDuty and New Relic
  • Serve as a senior escalation tier for high-severity incidents
  • Build runbooks for common failure modes and first-line resolution
  • Define and own SLIs and SLOs for critical services
  • Reduce MTTD and MTTR through instrumentation, alerting, diagnostics, and automation
  • Convert recurring support burden into permanent fixes, automation, or documentation
  • Maintain technology currency and lifecycle management for cloud services, Kubernetes clusters, operating systems, runtimes, and infrastructure components
  • Own business continuity and disaster recovery readiness, including backup and recovery strategies, recovery testing, failover capabilities, recovery runbooks, and RTO/RPO objectives
  • Lead infrastructure automation with Terraform
  • Eliminate toil through automation
  • Add reliability guardrails to CI/CD pipelines, including automated rollback, change-risk checks, and progressive delivery
  • Maintain production systems in accordance with regulatory and compliance requirements
  • Maintain business continuity and disaster recovery documentation and audit evidence
  • Partner with Security, Quality, and Compliance teams on audits, compliance, and remediation
  • Grow SRE and DevOps engineers through pairing, design and code review, and incident debriefs
  • Foster open communication among junior and contract engineers
  • Communicate documentation-first across distributed, multi-time-zone teams
  • Partner with software engineering, QA, and architecture to embed reliability into the development lifecycle
  • Inform capacity planning and scaling strategy with the Test team
  • Introduce proactive resilience testing such as game days
  • Align business continuity and disaster recovery capabilities with application requirements
  • Support cloud cost optimization through rightsizing, reserved capacity, and observability spend governance
  • Ensure work complies with company policies and applicable Privacy/HIPAA, regulatory, legal, and safety requirements

Requirements:

  • Demonstrated experience leading production support and incident management for production systems, including incident command during high-severity events
  • Strong grounding in SRE principles, including SLIs/SLOs, blameless postmortems, toil reduction, and reliability engineering
  • Demonstrated experience owning on-call strategy, including rotation design, alert tuning, and escalation
  • Expertise with Terraform or comparable IaC at scale, including module design, state management, and policy-as-code guardrails
  • Hands-on experience building CI/CD pipelines with reliability guardrails using GitHub Actions, Octopus Deploy, or Azure DevOps
  • Deep experience with at least one major cloud platform: AWS, Azure, or GCP
  • Experience with Docker and Kubernetes
  • Working knowledge of observability tooling such as Prometheus, Grafana, Datadog, CloudWatch, and ELK/OpenSearch
  • Experience designing and testing disaster recovery, including backup/restore, failover, and RTO/RPO validation
  • Working knowledge of cloud security and compliance practices, including IAM, network segmentation, encryption, vulnerability and patch management, and cloud cost optimization
  • Proficiency in at least one scripting or programming language such as Python, Go, or Bash
  • Experience in FDA and ISO regulated industries and agile methodologies preferred
  • B.S. in Computer Science or equivalent combination of education and applicable job experience, including technical school training and certifications in networks, servers, and cloud infrastructure; demonstrated production experience weighs more heavily than degree
  • Relevant cloud certifications such as AWS/Azure/GCP Professional or Architect level preferred
  • 10+ years in Site Reliability Engineering, DevOps, or infrastructure engineering
  • 2+ years mentoring or technically leading other engineers, including remote, offshore, or contracted partner engineers
  • Must be within the United States
  • Successful completion of pre-employment drug test and background check
  • Compliance with applicable company, Privacy/HIPAA, regulatory, legal, and safety requirements

Benefits:

  • Medical, dental, and vision benefits available the first day
  • Health savings accounts
  • Flexible savings accounts
  • 11 paid holidays per year
  • Minimum of 20 days of paid time off, with accrual starting on day 1
  • 401(k) plan with company match
  • Employee Stock Purchase plan
  • Equipment provided
  • Virtual training
  • Bonus and competitive compensation package
  • Joy-focused workplace supporting well-being, achievement, growth, fun, and camaraderie