Senior DevOps Engineer, Infrastructure – Reliability

Posted 11hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Senior DevOps Engineer improving Worth AI’s AWS infrastructure, Kubernetes reliability, and CI/CD automation. Building resilient, secure systems that make software delivery faster and easier for engineering teams.

Responsibilities:

  • Implement scalable Infrastructure-as-Code patterns using Terraform
  • Own and evolve the Kubernetes platform (EKS or self-managed), ensuring workloads are secure, scalable, and resilient
  • Optimize CI/CD pipelines to improve deployment frequency, reduce lead time, and increase release confidence
  • Design and enforce secure networking, IAM, and secrets management strategies across environments
  • Improve observability through metrics, logs, and tracing using tools such as DataDog
  • Optimize cloud cost efficiency through rightsizing, autoscaling, and architectural improvements
  • Implement disaster recovery planning, backup strategies, and multi-region resilience initiatives
  • Refactor brittle or manually managed infrastructure into automated, testable, reproducible systems
  • Introduce infrastructure tooling or architectural shifts and drive adoption through documentation, workshops, and hands-on support
  • Partner with engineering teams to eliminate friction in CI/CD, deployments, and cloud environments
  • Communicate technical trade-offs across engineering and product stakeholders
  • Maintain or exceed SLO/SLA targets, reduce incident frequency and duration, improve infrastructure stability and automation, and optimize cloud costs

Requirements:

  • 8+ years in DevOps, SRE, or infrastructure engineering
  • Proven experience designing and operating production Kubernetes environments at scale
  • Deep hands-on expertise with AWS infrastructure and cloud networking
  • Strong experience building and maintaining Terraform modules across large cloud environments
  • Demonstrated ownership of CI/CD systems and measurable improvement of DORA metrics
  • Experience leading incident response processes and driving meaningful postmortem outcomes
  • Strong understanding of distributed systems, event-driven architectures (Kafka), and database performance (PostgreSQL)
  • Proven ability to modernize legacy infrastructure and eliminate manual operational toil
  • Track record of taking a scoped infrastructure project from an ambiguous starting point to production without needing daily direction
  • Demonstrated ability to build trust across teams while raising the reliability bar
  • Bonus: Experience coding applications
  • Bonus: Experience operating high-throughput Kafka clusters (MSK or self-managed)
  • Bonus: Strong background in database performance tuning (PostgreSQL, Redis)
  • Bonus: Experience implementing autoscaling strategies for high-traffic systems
  • Bonus: Familiarity with service mesh technologies
  • Bonus: Experience building internal developer platforms (IDP)
  • Bonus: Background in security best practices (zero-trust networking, policy-as-code)
  • Bonus: Experience with multi-region or globally distributed systems
  • Bonus: Experience introducing platform-wide reliability frameworks (SLOs, error budgets, chaos testing)

Benefits:

  • Health Care Plan (Medical, Dental & Vision)
  • Retirement Plan (401k)
  • Life Insurance
  • Flexible Paid Time Off
  • 9 paid Holidays
  • Family Leave
  • Remote
  • Hybrid work (for Orlando Associates)
  • Free Food & Snacks (Orlando)
  • Wellness Resources