Site Reliability Engineer – Evening Shift
Posted 2hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Site Reliability Engineer maintaining reliable AWS and OpenShift production systems for Peraton, a national security and enterprise IT provider. Automating infrastructure, observability, deployments, and incident response.
Responsibilities:
- Operate and maintain production infrastructure services and applications, ensuring availability, reliability, performance, security, and operational health
- Monitor services and applications using SLIs, SLOs, dashboards, alerts, and observability tools
- Define application observability requirements and implement metrics, logs, traces, dashboards, and alerts
- Manage production incidents and service disruptions, including on-call response, troubleshooting, service restoration, root-cause analysis, and post-incident corrective actions
- Execute application and infrastructure releases through deployment pipelines, including staging and production promotion, validation, rollback, and release troubleshooting
- Manage the operational lifecycle of deployed infrastructure, including upgrades, patching, configuration changes, maintenance, and technology refreshes
- Assess and improve service resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing
- Identify and address reliability risks and operational technical debt using reliability metrics, incident trends, capacity data, and service health indicators
- Automate operational activities using an everything-as-code approach
- Collaborate with platform engineering and application teams to identify operational requirements and improve reliability and operability
Requirements:
- Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level clearance
- Bachelor's Degree and 8 years of experience, or a High School diploma/equivalent and 12 years of experience
- 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering
- Hands-on experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or comparable Kubernetes-based platforms
- Strong infrastructure-as-code experience with Terraform and Ansible/Ansible Tower
- Experience with CI/CD platforms GitLab and Jenkins, including reliability gating and deployment automation
- Proficient in Linux and Windows Server administration
- Experience with enterprise observability tools such as Dynatrace, Datadog, Splunk and Open Telemetry
- Demonstrated ownership of an SLI/SLO and alerting program, including error budgets, alert rationalization, and noise reduction
- Scripting/automation proficiency in Python, Bash, PowerShell, or Go
- Experience operating in federal or regulated environments (FISMA, FedRAMP, NIST 800-53)
- Preferred: AWS Solutions Architect, AWS DevOps Engineer, or AWS SysOps certification
- Preferred: Red Hat Certified Specialist in ROSA or Red Hat Certified System Administrator in OpenShift
- Preferred: Azure Administrator Associate or GCP Associate Cloud Engineer certification
- Preferred: Dynatrace Associate or Datadog Log Management Fundamentals certification
- Preferred: GitLab CI/CD Associate certification or Certified Jenkins Engineer (CJE)
- Preferred: Terraform Associate certification
Benefits:
- Overtime eligibility
- Shift differential
- Discretionary bonus
- Remote work









