Staff Observability Engineer
Posted 18hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Staff Observability Engineer designing AI observability and reliability systems for Tealium’s customer data platform. Building telemetry, SLOs, tracing, and cost controls across enterprise services.
Responsibilities:
- Lead end-to-end observability design, telemetry schemas, and OpenTelemetry pipeline architectures
- Partner with engineering and product teams to define SLIs, SLOs, error budgets, and production-readiness requirements
- Architect AI observability solutions for Amazon Bedrock and agentic workflows
- Track model selection, latency, retries, token usage, tool execution, and cost attribution
- Establish traceability across user requests, prompts, model calls, retrieved context, downstream services, and final responses
- Maintain data privacy and security controls
- Develop dashboards, alerts, SLOs, runbooks, and operational views
- Drive incident diagnosis, performance optimization, and economic sustainability
- Define measurable signals for response quality, groundedness, guardrail outcomes, and task completion
- Participate in an on-call rotation approximately 20% of the time
- Drive failure injection, production testing, and capacity planning
- Collaborate with SRE, MLOps, data engineering, security, product, and non-technical stakeholders
Requirements:
- 6+ years in Site Reliability, Observability, or Platform Engineering supporting 24x7x365 production systems
- Deep experience with OpenTelemetry and platforms such as Datadog, Sumo Logic, Prometheus, or Grafana
- Experience designing observability for distributed systems, APIs, event-driven architectures, and asynchronous workflows
- Hands-on experience instrumenting AI/ML or GenAI systems
- Proficiency with Amazon Bedrock or similar platforms
- Familiarity with agentic workflows, prompt engineering, vector databases, RAG architectures, LangChain, LlamaIndex, or SageMaker
- Proficiency in Java, Python, or Go
- Strong AWS expertise, including cloud networking, IAM, security, and service quotas
- Experience with IaC, CI/CD, and containers, including Terraform, Kubernetes, Argo CD, Jenkins, or GitHub Actions
- Experience with data modeling, telemetry pipelines, cardinality management, retention, responsible AI, and data privacy controls
- Strong communication, mentoring, and cross-functional leadership skills
Benefits:
- Performance-based bonus eligibility
- Equity options
- 15 hours of paid work time annually for volunteer activities and programs
- Remote-first working
- New hire stipends for home office setup
- New hire equity grants
- Paid time off
- Extended paid parental leave
- Company holidays
- Health and wellness programs covering physical, mental, social, and financial well-being
- Professional development opportunities
- Over 6,000 on-demand courses
- Manager and leadership development programs
- Health and related benefits programs
















