Cloud Infrastructure Engineer
Posted 1hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Cloud Infrastructure Engineer troubleshooting AWS and Azure production infrastructure for an automation-focused defense company. Automating operations with Linux, Python, Bash, Ansible, and AI tools.
Responsibilities:
- Serve as the first point of contact for cloud-related customer and application-team escalations
- Monitor production infrastructure and identify, isolate, and resolve issues before they impact business operations
- Execute automation scripts for cloud resource and OS-level operations
- Own assigned work, respond promptly to production issues, follow troubleshooting and escalation procedures, and maintain service availability and uptime
- Perform rapid triage and assess scope and business impact of cloud infrastructure problems
- Collect and analyze system logs, network traces, and filesystem information
- Perform basic Red Hat Linux LVM operations, analyze syslog, and troubleshoot OS-level processes and services
- Contact application owners or customers to gather missing information
- Isolate issues across application, operating system, and cloud infrastructure layers
- Open and manage support cases with AWS and Azure when needed
- Prepare logs, evidence, and preliminary troubleshooting for handoff to the senior L3 Infrastructure team when SLAs are exceeded
- Use AI assistants and AIOps tools for log analysis, issue identification, summarization, troubleshooting, and evidence gathering
- Monitor AWS and Azure Backup status reports and perform AWS AMI and Azure image-based VM restores
- Develop and maintain Bash, Python, and Ansible automation for data collection, filtering, troubleshooting, and evidence gathering
- Provide off-hours, night, or weekend support for critical infrastructure issues and escalations
Requirements:
- Must be a U.S. Citizen and only hold U.S. Citizenship; no dual citizens
- Security Clearance Level Required: Not Applicable
- 1–3 years of progressive IT experience, with hands-on exposure to cloud infrastructure administration and operational support
- Strong administration skills in Red Hat Linux
- Hands-on experience with AWS and Azure core services
- Proficiency in programming/scripting languages and automation tools such as Python, Bash, or Ansible
- Ability to safely execute automation scripts for cloud resource and OS-level operations, including understanding script inputs and outputs, workflow, and troubleshooting basic execution issues
- Familiarity with AI tools for troubleshooting, log analysis, and infrastructure support
- Willingness to actively troubleshoot issues, resolve problems in real time, and communicate directly with stakeholders
- Ability to follow established procedures, make appropriate decisions within defined guidelines, and recognize when an issue should be escalated to senior engineers
- Eagerness to learn from senior engineers and progressively take on more complex technical responsibilities
- Experience with AWS and Azure services including EC2, Virtual Machines, AMI, Managed Images, VPC/VNet, subnets, security groups/NSGs, VPN, Direct Connect, ExpressRoute, load balancers, S3, EBS, EFS, Azure Blobs, Managed Disks, Vault, IAM policies, RBAC, Secret Keys, and backup/restore operations
- Strong Red Hat Linux administration, including LVM/filesystem administration, system services and daemons, OS-level logging, performance tuning, and CPU/memory/IO troubleshooting
Benefits:
- 100% paid Medical, Dental & Vision for our employees
- 6% 401K match (Vested Immediately)
- 29 Days' PTO
- Flexible Work Schedule
- Tuition/Certification Reimbursement
- Growth Opportunities w/in an Emerging Defense Company


















