Principal Infrastructure Engineer

Posted 3ds ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Principal Infrastructure Engineer scaling Sezzle’s AWS, Kubernetes, and database platform. Improving reliability, performance, disaster recovery, and cost efficiency for its interest-free fintech payments business.

Responsibilities:

  • Design, build, operate, and scale Sezzle's infrastructure platform
  • Own complex technical initiatives from architecture and prototyping through implementation, production rollout, and ongoing operation
  • Identify system limits and implement improvements for increasing traffic, data volume, and workload complexity
  • Connect business workflows and transaction patterns to infrastructure improvements across applications, data, and systems
  • Build capacity models, run load and stress tests, diagnose bottlenecks, and validate throughput, latency, saturation, and cost improvements
  • Design resilient AWS account, IAM, network, and service architectures, including multi-AZ or multi-region solutions
  • Build and operate the Kubernetes platform, including lifecycle automation, workload isolation, resource allocation, autoscaling, upgrades, and deployment reliability
  • Scale and optimize Aurora RDS for MySQL and Postgres, including queries, indexes, connections, replication, failover, schema changes, and migrations
  • Define and instrument service-level objectives and error budgets; implement failure isolation, backpressure, load shedding, and safe retries
  • Participate in on-call and lead technical recovery during serious incidents and outages
  • Implement and test disaster recovery, backups, restores, and failover against recovery objectives
  • Build infrastructure-as-code and operational automation for provisioning, configuration, deployments, upgrades, and recovery
  • Improve observability through metrics, logs, traces, dashboards, and actionable alerts
  • Deliver safe infrastructure migrations with phased rollouts, validation, compatibility checks, and rollback paths
  • Improve cloud cost efficiency through resource right-sizing, utilization, autoscaling, storage tuning, and quantified savings
  • Build and evaluate AI-assisted tooling for incident investigation, runbooks, anomaly analysis, and toil reduction
  • Write architecture proposals, evaluate tradeoffs through prototypes and benchmarks, review shared-infrastructure changes, and document system operation and failure
  • Report to engineering leadership and collaborate with application engineering, Security, and Compliance

Requirements:

  • Bachelor's degree in Computer Science or a similar technical field (required)
  • 12+ years of experience across infrastructure, platform, site reliability, software development, or related engineering disciplines
  • Deep production expertise with AWS, including compute, IAM, multi-account architectures, VPC design, and private connectivity
  • Deep production expertise with Kubernetes, including cluster lifecycle, scheduling, resource management, autoscaling, networking, and troubleshooting
  • Deep expertise with RDS/Aurora MySQL and/or Postgres at scale
  • Track record of personally delivering infrastructure scaling improvements
  • Strong coding and automation skills using Golang, Python, or similar languages
  • Infrastructure-as-code experience with Terraform or equivalent
  • Strong systems fundamentals in Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes
  • Experience operating a 24/7 high-availability platform with direct customer or revenue impact
  • Hands-on incident response and postmortem remediation experience
  • Willingness to participate in an on-call rotation
  • Experience implementing and testing disaster recovery against defined recovery objectives
  • Practical experience with observability, load testing, capacity planning, and safe CI/CD practices
  • Active use of AI tooling in engineering or operations
  • Ability to carry ambiguous technical problems through production delivery and collaborate across engineering disciplines
  • EKS experience, fintech/payments/banking experience, multi-region architectures, chaos engineering, Prometheus/Grafana/Loki/Tempo, internal platform capabilities, and AI-assisted operational automation are preferred

Benefits:

  • Competitive gross monthly compensation of $12,500–$20,800 USD based on location and experience level
  • Open-source-focused technology environment