Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. As a Senior Site Reliability Engineer, you will lead observability and incident management efforts, ensuring system reliability and performance while mentoring engineers and enhancing operational processes.
Responsibilities:
- Own and continuously advance the observability strategy for your team — including monitoring, alerting, dashboards, distributed tracing, log aggregation, and SLI/SLO/SLA frameworks
- Lead incident management end-to-end: detection, triage, communication, resolution, and blameless post-mortems that drive lasting improvements
- Serve as a Reliability Engineering leader on your team — providing strong technical leadership, sound judgment, and a clear voice on reliability across the SDLC
- Design and maintain autonomous systems for building, deploying, testing, and operating all Filevine products with minimal human intervention
- Continuously enhance CI/CD pipelines, automation scripts, playbooks, and tooling to reduce toil and accelerate resolution time
- Proactively identify and resolve gaps in system availability, performance, and security while defending overall security posture
- Mentor engineers — dedicating meaningful time to growing those around you, sharing knowledge proactively, and leaving people more capable than before
- Document processes, architecture, procedures, and best practices; take full ownership of documentation for the technologies in your domain and actively close gaps for fellow SREs
- Participate in 24/7 on-call rotation for production support and emergency response; communicate clearly with technical and management stakeholders at all levels
- Build roadmaps for the technologies and workflows you own, giving the team a clear direction for ongoing improvement
Requirements:
- 8+ years of hands-on technical experience in software engineering, infrastructure, or operations roles, including a minimum of 5 years dedicated to Site Reliability Engineering
- Expert-level, well-rounded SRE skill set — proficient across monitoring/alerting, incident response, capacity planning, performance optimization, CI/CD, and reliability engineering best practices
- Deep hands-on expertise with New Relic or a comparable observability platform; strong preference for candidates who have led observability platform adoption or migration at scale
- Demonstrated experience owning incident management programs: on-call processes, escalation design, post-mortem culture, and measurable MTTR/MTTD improvement
- Strong proficiency in Python, Bash, PowerShell, and other common SRE scripting and automation technologies
- Expert-level experience designing, building, and maintaining autonomous systems that handle software build, deployment, testing, monitoring, and operations
- Proficient hands-on experience with AWS (EC2, EKS/Kubernetes, CloudWatch, Lambda, S3, IAM) and the broader cloud-native ecosystem
- Strong communicator who proactively informs stakeholders, operates transparently, and can bridge technical complexity for product and management audiences
- Proven track record of mentoring engineers, leading initiatives to completion, and making those around them measurably better
- Bachelor's degree in Computer Science, Information Systems, or a related field; equivalent certifications (e.g., AWS certifications, Google Cloud Professional); or substantial comparable direct work experience