Job Title: Site Reliability & Operations Engineer (L3 Support)
Location: Blue Ash, OH (5 days onsite)
Job Duration: 18 Month(s) - CTH
Travel: Required (approximately every 4 weeks; additional travel during new site go-lives)
Key Responsibilities:
- Lead and manage Major Incident (P1/P2) response, including bridge/war-room coordination and stakeholder communication.
- Drive Root Cause Analysis (RCA), Problem Management, and corrective actions through to closure.
- Provide L3 production support for business-critical applications and infrastructure.
- Monitor, troubleshoot, and resolve production issues across cloud, on-premises, and distributed environments.
- Partner with development, infrastructure, and business teams to resolve incidents and improve platform reliability.
- Participate in application design, testing, deployment, and software delivery improvements.
- Support infrastructure and applications through a rotating 24x7 on-call schedule.
- Create and maintain operational procedures, documentation, and support processes.
Required Skills:
- Hands-on experience leading Major Incident Management (P1/P2) from start to resolution.
- Strong experience with Root Cause Analysis (RCA) and Problem Management methodologies (5 Whys, Fishbone, Timeline Analysis, etc.).
- Experience providing L3 Production Support or Site Reliability Engineering (SRE) support in enterprise environments.
- Strong troubleshooting skills using monitoring tools such as Dynatrace, Splunk, Grafana, AppDynamics, or similar.
- Experience supporting distributed applications across cloud and on-premises environments.
- Excellent communication and stakeholder management skills.