evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk
design and implement processes that enforce enterprise resiliency and reliability standards
lead blameless post‑incident reviews for high‑severity incidents
partner with product and platform teams to proactively identify and remediate reliability risks
develop, communicate, and evangelize new standards, tools, and frameworks across subdivisions
troubleshoot complex production issues and implement durable solutions
participate in a periodic on‑call rotation to support production stability
evaluate and onboard resiliency and reliability tooling
actively participate in reliability engineering and resilience communities of practice
Requirements
experience with modern observability and monitoring tools, such as Splunk, Honeycomb, CloudWatch, Dynatrace, or AppDynamics
strong understanding of SLIs, SLOs, and SLAs
experience with alert design, anomaly detection, predictive alerting, and synthetic monitoring
experience with automation and resilience practices such as Python-based automation, RPA platforms (e.g., Blue Prism, UiPath), chaos engineering, and failure analysis techniques (e.g., FMEA)