DevOps Engineer IV serves as a senior individual contributor and technical leader within a team, providing leadership, guidance, and mentoring to other engineers.
You will be responsible for meeting scope, schedule, and delivery requirements, interacting with stakeholders, and driving improvements in DevOps processes and practices across the program.
Set up and operate monitoring and observability tooling (AWS CloudWatch, Prometheus, Grafana, Loki log aggregation) for real-time visibility into application health, performance, and infrastructure
Build and maintain "Golden Signals" performance dashboards measuring latency, traffic, errors, and saturation
Implement the DORA metrics roadmap using Grafana and GitLab analytics to establish performance baselines
Manage Tier 2/3 production support within strict SLAs: 1-hour initial response, 4-hour critical resolution, 99.9% uptime commitment
Author Root Cause Analyses within 3 business days of any severity-1 production outage; maintain on-call runbooks and change correlation
Author and maintain the BCDR plan, including recovery architecture and RTO targets, cross-region replication (RDS, S3), Route 53 routing, and Secrets Manager; coordinate biannual failover drills
Configure centralized alerting and incident tooling (Jira Service Desk/ServiceNow, Microsoft Teams, AWS Chatbot)
Implement AWS Auto Scaling and Elastic Load Balancing; deliver sprint performance reports and cost-optimization recommendations
Support recruiting efforts by evaluating homework assignments and potentially assisting with interviews
Requirements
Bachelor's degree and 8+ years of relevant experience, or equivalent additional experience in lieu of a degree
Must meet federal suitability requirements and pass a background investigation as a condition of employment
5+ years of hands-on experience with AWS, Terraform (or similar IaC), and Git/GitLab in production environments
Experience supporting 5 or more engineering teams from a shared DevOps/platform function
Demonstrated experience designing and building CI/CD pipelines in GitLab and/or Jenkins, including quality and security gates
Strong working knowledge of containerization using Docker and orchestration on AWS ECS/EKS/Fargate
Familiarity with DevSecOps practices including SAST, dependency scanning, and automated vulnerability remediation
Excellent communication and documentation skills; comfortable in a highly collaborative Agile/SAFe environment
Proven experience in SRE, production operations, or incident response for mission-critical, high-availability cloud-native services
BCDR planning and disaster recovery exercise experience
Tech Stack
AWS
Cloud
Docker
Grafana
Jenkins
Prometheus
ServiceNow
Terraform
Benefits
Company-subsidized health, dental, and vision insurance