Senior AIOps Platform Engineer
Posted 3ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior AIOps Strategy & Implementation Engineer at EY optimizing observability platforms and building automation workflows in cloud and on-prem environments.
Responsibilities:
- Implement and integrate AIOps solutions across cloud and on-prem environments to enhance incident detection, correlation, and automated remediation.
- Configure, manage, and optimize observability platforms (logs, metrics, traces) including ingestion pipelines, dashboards, alerting rules, and topology views.
- Build automation workflows for operational tasks such as event triage, remediation actions, resource optimization, and predictive maintenance.
- Collaborate with platform, cloud, and operations teams to onboard new services, data sources, and operational signals into the AIOps ecosystem.
- Support deployment, integration, and operationalization of AI/ML models used for anomaly detection, event correlation, forecasting, and noise reduction (not model training).
- Develop and implement playbooks, runbooks, and automated actions in line with SRE, DevOps, and cloud operations practices.
- Troubleshoot issues related to observability ingestion, integration connectors, pipeline performance, and automation workflows.
- Participate in architectural discussions, tool evaluations, and AIOps capability uplift initiatives.
- Ensure compliance with security, governance, IAM, and data handling guidelines in the AIOps and observability ecosystem.
- Document configurations, workflows, integrations, and operational guidelines for AIOps platform usage.
Requirements:
- Up to 8 years of experience in cloud/infrastructure operations, SRE, observability engineering, or automation engineering.
- Strong hands-on experience with observability tools (e.g., ELK/Elastic, Prometheus, Grafana, Dynatrace, AppDynamics, Datadog, New Relic, Splunk, or equivalent).
- Experience building workflows using automation frameworks or orchestration tools (e.g., Ansible, ServiceNow, Rundeck, Azure Automation, Lambda/Functions).
- Familiarity with AI/ML model deployment concepts for IT operations (consumption of models, not development).
- Understanding of event management, correlation engines, topology mapping, and CMDB integrations.
- Proficiency with scripting languages (Python, PowerShell, Bash) for automations and integrations.
- Knowledge of cloud platforms (Azure/AWS/GCP), their monitoring services, and operational data flows.
- Bachelor’s degree in Information Technology, Computer Science, or related discipline.
Benefits:
- Competitive salary



















