Senior Site Reliability Engineer
Posted 5hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior SRE operating production systems and building LLM agents for Mozn, an enterprise AI company. Automating incident response, reliability toil, and observability workflows with guardrails.
Responsibilities:
- Carry a normal on-call rotation and act as a hands-on incident responder
- Investigate, fix, and document production incidents
- Perform application-level debugging and identify root causes in service code and business logic
- Ship fixes or pull requests directly into application repositories when appropriate
- Design, build, and ship LLM-based agents integrated with Kubernetes, cloud APIs, observability, and incident-management tools
- Define agent tool interfaces and build safe wrappers for APIs, scripts, and read/write actions
- Establish autonomy guardrails and human-approval requirements for agents
- Own agent evaluations and build test/backtest suites using historical incidents
- Tune prompts, context, and tool schemas as agent scope expands
- Partner with the SRE/platform team to identify suitable automation workflows
- Report agent impact using MTTD, MTTR, MTTX, false-positive/negative rates, and engineer-hours of toil removed
- Maintain security- and compliance-first operations with audit trails, least-privilege production access, and Saudi data-residency/regulatory alignment
Requirements:
- 3+ years building production software with LLMs, including agentic workflows, tool/function calling, multi-step planning, or RAG
- Hands-on experience shipping work with Claude Code, OpenAI Codex, or Kimi K2/K3
- Strong Python or similar programming skills for agent tooling, API wrappers, and orchestration
- Hands-on SRE experience as a primary on-call responder, including incident response and root cause analysis
- Application-level debugging skills and ability to read service code, trace failures to underlying logic, and ship fixes
- Hands-on Kubernetes and cloud provider experience with AWS, GCP, OCI, or Azure
- Fluency with Prometheus, Grafana, Datadog, or ELK
- Understanding of autonomous-system guardrails, permissioning, approval gates, rollback paths, and auditability
- Ability to build trust with technical stakeholders
- Experience in Saudi Arabia/MENA, Terraform/Ansible, Docker, VM/on-prem setups, LLM agent evaluation, ML engineering, LLMOps, or platform engineering are nice to have
Benefits:
- Competitive compensation
- Top-tier health insurance
- Responsibility and trust
- Freedom and autonomy in decision-making
- Fun and dynamic workplace
- Opportunity to work alongside leading AI talent
- Inclusive and empowering culture
- Opportunity to work at the forefront of AI in the Middle East
















