Senior Site Reliability Engineer
San Francisco, CA, USA
- Deep SRE Expertise. 6+ years of hands-on SRE, DevOps, or Production Engineering experience in high-scale digital applications, with a strong understanding of reliability principles and operational excellence.
- Cloud-Native Technical Skills. Strong exposure to Azure AKS, Kubernetes, Docker, Service Mesh, and API-driven architectures, with operational support experience for React front-end and Spring Boot microservices in production environments.
- Observability and Automation Mastery. Hands-on experience with observability tools (Dynatrace, Splunk, Grafana, Prometheus) and strong scripting abilities (Python, Bash, PowerShell, YAML) to build automation that reduces toil and improves incident response.
- Incident Management Excellence. Proven experience in incident management, root cause analysis, and implementing permanent corrective actions that drive long-term reliability improvements.
- CI/CD and Platform Knowledge. Experience with SRE principles, CI/CD pipelines (Jenkins, GitHub Actions), and cloud platforms (Azure required; AWS/Google Cloud Platform/OCI a plus).
- Analytical Problem-Solver. Strong analytical and problem-solving abilities with clear communication skills under pressure, a collaborative mindset, and passion for reducing toil while improving developer and operator experiences.