Oracle is a leading company in AI and cloud solutions, and they are seeking a Senior Site Reliability Engineer to join their OCI team. This role involves minimizing downtime of OCI services by managing incidents and enhancing service availability through automation and operational excellence.
Responsibilities:
- Solve complex problems related to infrastructure cloud services and automate common tasks to ensure continuous availability with minimal human intervention
- Command and coordinate SMEs and service leaders to restore services as quickly as possible during major incidents, while keeping accurate and timely data on the progress of such incidents
- Utilize a deep understanding of cloud computing design patterns and their dependencies to mitigate complex major incidents
- Embed a methodical approach to troubleshoot large, complex, interconnected systems used in incident detection and orchestration
- Document pertinent information related to incidents that aids process improvement, identifies deviations, and enables the creation of an incident knowledge base
- Monitor and evaluate high-level service and infrastructure dashboards, taking action to address identified anomalies
- Identify opportunities and take ownership of automation and/or continuous improvement of incident management process steps and best practices
- Define and document the technical architecture of large-scale distributed systems
- Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services
- Be responsible for the design and delivery of the mission-critical stack, with a focus on security, resiliency, scalability, and performance
- Partner with development teams to define operational requirements for product roadmaps
- Articulate the technical characteristics of services and technology areas, and guide development teams to engineer and add premier capabilities to the Oracle Cloud service portfolio
- Act as the ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs)
Requirements:
- Bachelor's degree or higher in Computer Science or relevant work experience
- 3+ years' experience in Site Reliability Engineering, DevOps, or System Engineering
- Must have public cloud operations experience (e.g., AWS, Azure, GCP, OCI)
- Extensive experience with Major Incident Management in a cloud-based environment
- Demonstrate clear understanding of automation and orchestration principles
- Experience having worked in at least one modern object-oriented programming language
- Experience with professional software engineering standard methodologies such as Agile project management, coding standards, code reviews, source control management, build processes, testing, and operations
- Familiarity with infrastructure automation tools such as Chef, Ansible, Jenkins, Terraform
- Excellent expertise with several of following technologies: Infrastructure-as-a-Service, CI/CD systems, Docker, RESTful APIs, log analysis tools, debugging tools
- 8 years of experience in software engineering, infrastructure management, or related field
- OR Bachelor's Degree in Computer Science, Engineering, or related field AND 4 years of experience in software engineering, infrastructure management, or related field
- OR Master's Degree in Computer Science, Engineering, or related field AND 2 year of experience in software engineering, infrastructure management, or related field
- OR Doctorate in Computer Science, Engineering, or related field
- 3 years of experience in automation
- 3 years of experience in programming and/or scripting
- 9 years of experience in software engineering, infrastructure management, or related field
- Bachelor's Degree in Computer Science, Engineering, or related field AND 5 years of experience in software engineering, infrastructure management, or related field
- Master's Degree in Computer Science, Engineering, or related field AND 3 years of experience in software engineering, infrastructure management, or related field
- Doctorate in Computer Science, Engineering, or related field AND 1 year of experience in software engineering, infrastructure management, or related field
- 5 years of experience in automation
- 5 years of experience in programming and/or scripting