NextAmp LLC is a digital modernization services company focused on transforming the insurance industry. They are seeking a highly skilled DevOps / Site Reliability Engineer (SRE) to build, automate, and maintain reliable, scalable, and secure cloud infrastructure.
Responsibilities:
- Design, implement, and maintain CI/CD pipelines to enable reliable and automated software deployments
- Build, manage, and optimize containerized applications using Docker and Kubernetes
- Provision and manage cloud infrastructure using Infrastructure as Code (IaC) tools such as Terraform or CloudFormation
- Deploy, monitor, and maintain applications on AWS or Azure
- Ensure high availability, scalability, security, and reliability of production environments
- Monitor application and infrastructure health, respond to incidents, and perform root cause analysis (RCA)
- Automate operational tasks to improve efficiency and reduce manual effort
- Collaborate with development teams to improve deployment processes and application reliability
- Implement monitoring, logging, alerting, and observability best practices
Requirements:
- 4+ years of experience in DevOps or Site Reliability Engineering (SRE)
- Strong experience building and managing CI/CD pipelines using tools such as Jenkins, GitHub Actions, Azure DevOps, or GitLab CI
- Hands-on experience with Docker and Kubernetes
- Strong knowledge of Infrastructure as Code (Terraform, CloudFormation, or similar)
- Experience with AWS or Azure cloud platforms
- Experience with monitoring and observability tools such as Prometheus, Grafana, CloudWatch, Datadog, Splunk, ELK, or Azure Monitor
- Good understanding of Linux system administration, networking, and security best practices
- Experience with scripting using Bash, Python, or PowerShell
- Strong troubleshooting and production incident management skills
- Experience with Helm, ArgoCD, or FluxCD
- Knowledge of container security and vulnerability management
- Familiarity with service mesh technologies (Istio, Linkerd)
- Experience with secrets management tools such as HashiCorp Vault or AWS Secrets Manager
- Knowledge of high availability, disaster recovery, and backup strategies