NetSuite is a subsidiary of Oracle that provides cloud-based business management software. As a Senior Site Reliability Engineer, you will play a pivotal role in building and operating the Oracle HealthPatient Portal, focusing on designing and operating scalable infrastructure while improving system reliability through automation and observability practices.
Responsibilities:
- Work with the Site Reliability Engineering (SRE) team to take shared ownership of services and platform components. Develop a strong understanding of end-to-end system architecture, dependencies, and production behavior
- Design, build, and operate reliable, scalable, and secure infrastructure supporting large-scale distributed systems
- Improve system reliability through automation, monitoring, and performance optimization
- Contribute to the adoption of AI-assisted approaches for operations, including: Enhancing observability and alerting Supporting automated incident detection and remediation Exploring intelligent automation for infrastructure lifecycle management
- Partner with development teams to enhance service architecture, scalability, and operability
- Participate in on-call rotations and act as an escalation point for complex production issues
- Perform root cause analysis and implement long-term fixes to prevent recurrence
- Apply knowledge of distributed systems to troubleshoot issues and optimize system performance
- Drive continuous improvement in DevOps/SRE practices, including CI/CD, Infrastructure as Code, and automation at scale
Requirements:
- Experience building and operating high-availability, fault-tolerant systems
- Strong understanding of distributed systems, performance monitoring, and resiliency patterns
- Experience with incident response, root-cause analysis, and production troubleshooting
- Experience with one or more cloud environments OCI, AWS/Azure
- Advanced competency in CI/CD pipelines (Jenkins, Kubernetes)
- Infrastructure as Code (Terraform)
- Observability tools (Prometheus, Grafana)
- Strong focus on automation-first operations
- Proficiency in Data Warehousing platforms (e.g., Vertica, Snowflake)
- Experience with ETL frameworks and large-scale data processing
- Understanding of columnar storage systems
- Proficiency in Python, Java, or Go
- Experience with Docker, Kubernetes, and shell scripting
- Strong troubleshooting skills with ability to perform root-cause analysis
- Experience resolving complex production issues in distributed systems
- Apply DevOps/SRE practices to automate deployments and operations
- Enhance observability using Prometheus/Grafana and AI-driven insights
- Participate in on-call rotations
- Implement preventative and automated remediation solutions
- Work closely with engineers to execute technical roadmaps
- Contribute to code reviews and infrastructure improvements
- 4+ years of software engineering, cloud infrastructure, SRE, or DevOps experience
- Proven ownership of production system reliability in cloud environments
- Cloud infrastructure design and automation
- Distributed systems and performance optimization
- Data warehousing and ETL frameworks
- Terraform, Docker, Kubernetes
- Observability stacks (Prometheus, Grafana)
- Python, Java, or Go
- Strong problem-solving mindset with a focus on automation and scalability
- Experience improving system reliability through intelligent automation
- Experience in healthcare or regulated environments (HIPAA, compliance frameworks)
- Experience working in environments requiring security clearance
- Experience building self-healing or autonomous infrastructure systems