Infrastructure & Platform Operations Engineer

Posted 11hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Infrastructure & Platform Operations Engineer supporting DysrupIT’s Nephos technology services. Operating AWS, Linux, databases, containers, monitoring and enterprise production platforms.

Responsibilities:

  • Support, maintain and improve infrastructure and production platforms used to deliver Nephos services to customers
  • Provide BAU production support, service stability and knowledge transfer across cloud and on-premise infrastructure
  • Monitor infrastructure and applications, investigate alerts and logs, identify root causes and reduce unnecessary alert noise
  • Administer and troubleshoot customer environments, Linux systems, networking, PostgreSQL, MongoDB, Docker, Kubernetes and AWS container services
  • Use Python, PowerShell and other scripting and automation to improve repeatability and reduce manual effort
  • Manage incidents, problems, service requests and changes, including troubleshooting, service restoration, root-cause analysis, validation and rollback planning
  • Support enterprise platform installations, configuration, upgrades, testing, validation and recovery
  • Create and maintain knowledge articles, runbooks and work instructions, and participate in cross-training
  • Work with customer and internal technical teams during incidents, changes, upgrades and investigations
  • Collaborate with DevOps, Data Services, Product, Delivery and other cross-functional teams
  • Contribute to platform improvements, resilience, capacity, monitoring and operational design
  • Support backup, restore, disaster-recovery activities and practical recovery exercises
  • Apply security principles including least privilege, secure credential handling, secrets, certificates, auditability, patching and vulnerability awareness

Requirements:

  • Typically 5+ years of experience in infrastructure engineering, platform operations, systems engineering or production support; demonstrable capability is more important than a fixed number of years
  • Strong BAU and production support experience, including live incidents, service restoration, monitoring, failed changes or deployments, root-cause analysis and controlled remediation
  • Strong hands-on AWS production experience; AWS is the primary cloud requirement
  • Very strong Linux administration and troubleshooting experience, particularly Ubuntu and Red Hat, including command-line use
  • Deep networking and connectivity troubleshooting experience covering TCP/IP, DNS, routing, firewalls, VPNs, TLS and cloud networking
  • Strong production MongoDB administration and troubleshooting experience
  • Meaningful PostgreSQL operational experience, including administration, backup and recovery, monitoring and troubleshooting
  • Hands-on Docker production experience and strong knowledge of container lifecycle, networking, logging, health and troubleshooting
  • Strong hands-on Kubernetes experience, including deployment health, pods, services, configuration, logs and failure diagnosis
  • Experience supporting enterprise applications and platforms through installation, configuration, upgrades, health monitoring and technical troubleshooting
  • Strong scripting and automation capability, particularly Python and/or PowerShell; Bash or other transferable scripting languages relevant
  • Strong monitoring and observability fundamentals, including infrastructure metrics, log investigation and alert analysis
  • Working knowledge of Git and GitHub
  • Practical understanding of backup, restore, recovery and service resilience
  • Practical IT service management experience covering Incident, Problem, Change, Service Request and major-incident processes
  • Experience with a service-management platform such as Jira/JSM, ServiceNow, Remedy, Freshservice or equivalent
  • Good security awareness, including least privilege, IAM concepts, secrets, certificates, MFA, privileged access, patching and vulnerability awareness
  • Strong customer-centric attitude and clear communication with technical and non-technical stakeholders
  • Excellent written and verbal communication skills, including technical documentation and operational procedures
  • Ability to learn complex unfamiliar technology and become independently effective following onboarding and knowledge transfer
  • Ability to work independently, manage multiple priorities and recognise when escalation or wider technical input is required
  • Willingness to participate in occasional out-of-hours upgrades, changes and major incidents
  • Nice-to-have: Azure, Google Cloud Platform, infrastructure as code, CI/CD, BigID, RabbitMQ, Redis, LogicMonitor, privileged-access and secrets-management technologies, REST APIs, TLS certificates, disaster-recovery planning, data governance, relevant certifications, microservices, serverless or distributed architectures

Benefits:

  • Government-mandated benefits plus supplemental HMO coverage
  • Collaborative and professional work environment
  • Career growth opportunities within a growing technology organization
  • Hybrid/Remote work arrangement (where applicable)
  • Competitive compensation package commensurate with experience