Infrastructure & Platform Operations Engineer
Posted 11hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Infrastructure & Platform Operations Engineer supporting DysrupIT’s Nephos technology services. Operating AWS, Linux, databases, containers, monitoring and enterprise production platforms.
Responsibilities:
- Support, maintain and improve infrastructure and production platforms used to deliver Nephos services to customers
- Provide BAU production support, service stability and knowledge transfer across cloud and on-premise infrastructure
- Monitor infrastructure and applications, investigate alerts and logs, identify root causes and reduce unnecessary alert noise
- Administer and troubleshoot customer environments, Linux systems, networking, PostgreSQL, MongoDB, Docker, Kubernetes and AWS container services
- Use Python, PowerShell and other scripting and automation to improve repeatability and reduce manual effort
- Manage incidents, problems, service requests and changes, including troubleshooting, service restoration, root-cause analysis, validation and rollback planning
- Support enterprise platform installations, configuration, upgrades, testing, validation and recovery
- Create and maintain knowledge articles, runbooks and work instructions, and participate in cross-training
- Work with customer and internal technical teams during incidents, changes, upgrades and investigations
- Collaborate with DevOps, Data Services, Product, Delivery and other cross-functional teams
- Contribute to platform improvements, resilience, capacity, monitoring and operational design
- Support backup, restore, disaster-recovery activities and practical recovery exercises
- Apply security principles including least privilege, secure credential handling, secrets, certificates, auditability, patching and vulnerability awareness
Requirements:
- Typically 5+ years of experience in infrastructure engineering, platform operations, systems engineering or production support; demonstrable capability is more important than a fixed number of years
- Strong BAU and production support experience, including live incidents, service restoration, monitoring, failed changes or deployments, root-cause analysis and controlled remediation
- Strong hands-on AWS production experience; AWS is the primary cloud requirement
- Very strong Linux administration and troubleshooting experience, particularly Ubuntu and Red Hat, including command-line use
- Deep networking and connectivity troubleshooting experience covering TCP/IP, DNS, routing, firewalls, VPNs, TLS and cloud networking
- Strong production MongoDB administration and troubleshooting experience
- Meaningful PostgreSQL operational experience, including administration, backup and recovery, monitoring and troubleshooting
- Hands-on Docker production experience and strong knowledge of container lifecycle, networking, logging, health and troubleshooting
- Strong hands-on Kubernetes experience, including deployment health, pods, services, configuration, logs and failure diagnosis
- Experience supporting enterprise applications and platforms through installation, configuration, upgrades, health monitoring and technical troubleshooting
- Strong scripting and automation capability, particularly Python and/or PowerShell; Bash or other transferable scripting languages relevant
- Strong monitoring and observability fundamentals, including infrastructure metrics, log investigation and alert analysis
- Working knowledge of Git and GitHub
- Practical understanding of backup, restore, recovery and service resilience
- Practical IT service management experience covering Incident, Problem, Change, Service Request and major-incident processes
- Experience with a service-management platform such as Jira/JSM, ServiceNow, Remedy, Freshservice or equivalent
- Good security awareness, including least privilege, IAM concepts, secrets, certificates, MFA, privileged access, patching and vulnerability awareness
- Strong customer-centric attitude and clear communication with technical and non-technical stakeholders
- Excellent written and verbal communication skills, including technical documentation and operational procedures
- Ability to learn complex unfamiliar technology and become independently effective following onboarding and knowledge transfer
- Ability to work independently, manage multiple priorities and recognise when escalation or wider technical input is required
- Willingness to participate in occasional out-of-hours upgrades, changes and major incidents
- Nice-to-have: Azure, Google Cloud Platform, infrastructure as code, CI/CD, BigID, RabbitMQ, Redis, LogicMonitor, privileged-access and secrets-management technologies, REST APIs, TLS certificates, disaster-recovery planning, data governance, relevant certifications, microservices, serverless or distributed architectures
Benefits:
- Government-mandated benefits plus supplemental HMO coverage
- Collaborative and professional work environment
- Career growth opportunities within a growing technology organization
- Hybrid/Remote work arrangement (where applicable)
- Competitive compensation package commensurate with experience

















