Senior Observability Engineer
Posted 1hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Sr. Observability Engineer II improving Dynatrace monitoring and alert quality for NextGen Healthcare’s client-facing healthcare SaaS environment. Reducing operational noise and strengthening service reliability.
Responsibilities:
- Serve as the subject matter expert for the Dynatrace observability platform
- Mature proactive monitoring capabilities through signal quality engineering and operational noise reduction
- Act as a technical liaison between Hosting Operations, Site Reliability Engineering, and other engineering teams
- Engineer and maintain the Dynatrace multi-tenant environment configuration, including management zones, alerting profiles, metric and custom events, tagging strategies, and access boundaries
- Own the end-to-end alert lifecycle, including ownership, naming standards, severity mapping, alert disposition taxonomies, and production-readiness criteria
- Eliminate alert fatigue through tuning, thresholds, correlation, suppression, and alert retirement while maintaining coverage for client-impacting issues
- Design and implement synthetic monitors and business-centric service-health models for critical client journeys and access paths
- Author and standardize Dynatrace Query Language queries, notebooks, and dashboards for operations, leadership, and reliability reporting
- Manage observability configurations using configuration-as-code practices so monitoring setups are versioned, peer-reviewed, and reproducible
- Enhance alert-to-ticket enrichment, including Salesforce ITSM integrations, with context and probable cause
- Partner with SQL, platform, OS, and API specialists to resolve monitoring blind spots and ensure alerts have robust runbooks
- Establish reporting baselines for P1–P4 service-level measurements and pair noise-reduction changes with operational guardrails
- Maintain reference materials, standard operating procedures, and observability standards documentation in Confluence
- Perform other duties supporting the overall objective of the position
Requirements:
- Bachelor’s Degree in Computer Science, Information Technology, or a related technical field, or any combination of education and experience providing the required qualifications
- 7+ years of experience in observability, application performance monitoring (APM), monitoring engineering, site reliability engineering (SRE), or closely related technology infrastructure
- Hands-on expertise with the Dynatrace platform, including Davis AI, DQL, dashboards, notebooks, management zones, alerting profiles, metric/custom events, and synthetic monitors
- Experience designing, configuring, and tuning enterprise-scale alerting systems to reduce noise while maintaining detection capabilities
- Experience integrating monitoring systems with enterprise ITSM/ticketing systems such as Salesforce for automated ticket routing and enrichment
- Experience working within highly regulated hosting environments, such as healthcare, HIPAA/HITRUST, SOC 2, or ISO 27001
- Dynatrace certification, or commitment to obtain certification within 1 year of hire
- Strong working knowledge of AWS infrastructure, Windows and Linux operating systems, and SQL Server database concepts
- Advanced understanding of cloud-native observability frameworks, APM tools, specifically Dynatrace, and alert lifecycle management
- Strong command of AWS cloud computing, systems integration, database operations, and scripting such as Python or PowerShell
- Excellent problem-solving capabilities
- Strong written and verbal communication skills
- Highly organized, detail-oriented, and self-driven, with a strong ownership mindset
- Ability to translate complex technical infrastructure metrics into direct service and business impacts
- Ability to collaborate across operations command analysts, SRE partners, database specialists, and leadership
Benefits:
- Equal opportunity employer
- Inclusive and diverse work environment

















