AI Pipeline Engineer – Security Automation Platform

Posted 1hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

AI Pipeline Engineer building reliable threat-intelligence automation for CloudLinux's Imunify360 security platform. Operating large-scale releases, quality gates, observability and rollback systems.

Responsibilities:

  • Design, build and operate automated pipelines end to end
  • Turn fragile multi-stage batch jobs into resumable, idempotent and observable systems with explicit state machines and recovery paths
  • Define and enforce latency budgets and SLOs per stage, making violations visible and actionable
  • Build observability layers with metrics, dashboards, alerting and health gates
  • Design and implement automatic holds, rollbacks, blast-radius limits, kill switches and safe-by-default behavior
  • Eliminate manual steps and reduce operational maintenance
  • Write and maintain unit and integration tests for concurrency, partial failure, external API flakiness and multi-stage state
  • Investigate and resolve issues across ClickHouse, GitLab CI, S3/object storage, Prometheus/Grafana and third-party APIs
  • Collaborate with security analysts and the Server team on architecture and production-ready designs
  • Own automated protection, progressive release, quality gates, CI at scale, LLM orchestration and observability for Imunify360's security platform

Requirements:

  • 5+ years of professional backend, platform or infrastructure engineering experience
  • Demonstrable experience building and operating multi-stage data or automation pipelines
  • Real depth in at least one of Python, Go or Rust
  • Systems design judgement and experience designing reliable production systems
  • Practical experience with workflow orchestration and job scheduling
  • Reliability engineering knowledge including idempotency, retries with backoff, checkpointing, resumability, graceful degradation, backpressure and partial-failure handling
  • Hands-on observability experience with Prometheus/Grafana, LGTM stack or equivalent, including designing metrics
  • Deep CI/CD experience, ideally GitLab CI with dynamic/child pipelines and self-hosted runners
  • Comfort with Docker and container-based test environments
  • Experience with S3/Ceph or equivalent object storage
  • Experience with ClickHouse or another columnar database
  • Comfort designing state machines and long-running processes that survive restarts
  • Ability to reason about concurrency across multiple in-flight rollouts
  • Excellent debugging skills across system, network and data layers
  • Strong communication skills and comfort working in a distributed team
  • At least upper-intermediate proficiency in spoken and written English

Benefits:

  • A strong focus on professional development with opportunities for learning and growth: interesting and challenging projects, mentor and other knowledge-exchange programs
  • Fully remote work with flexible working hours
  • Paid 24 days of vacation per year
  • 10 days of national holidays
  • Unlimited sick leaves
  • Compensation for private medical insurance
  • Co-working reimbursement
  • Gym/sports reimbursement
  • Opportunity to receive a reward for the most innovative idea that the company can patent