Cribl is building the telemetry infrastructure for the AI era, partnering with IT and Security teams at major enterprises. They are seeking a Senior Site Reliability Engineer to enhance observability, reliability, and operational excellence across their products.
Responsibilities:
- Engage with teams and improve service delivery and reliability across their entire lifecycle
- Measure and monitor all production systems with an eye towards availability, latency and overall system health
- Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence
- Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability
- Help identify and drive down toil with creative innovation and automation
- This position will require stand-by, on-call, or off-hours duties
Requirements:
- Proven experience designing, implementing, and operating observability systems for complex cloud-based platforms, with deep knowledge of best practices and a strong drive to implement them leveraging Cribl products
- Knowledge of cloud platforms (prefer AWS and Azure) and container + orchestration technologies
- Experience with APM and Observability and related tools such as, New Relic, Splunk, CloudWatch, Prometheus, Grafana/Kibana, Sentry etc
- Extensive experience with enterprise scale continuous delivery environments
- Development with JavaScript/Node.js/TypeScript in a Linux/Mac environment
- Experience with sustainable incident response in a blameless environment
- Background in Linux Systems Engineering
- Experience with Incident response related tools for instance, PagerDuty, FireHydrant, Blameless etc
- Comfortable with a high level of autonomy and working with a distributed team
- Knowledge of Cloud and application security best practices
- Strong knowledge of cloud design patterns for scale, data management, resiliency, etc
- A love for high quality and a knack for testing
- Opinions about business metrics, and SLOs
- Experience with Configuration Management and Infrastructure as a Code Tools like Terraform (preferred) or Ansible
- Experience working with Cloud SDKs is also a plus