Job Title: AMS Lead (Application Management Services) SRE / Incident Management
Location: Remote / Hybrid / Onsite
Job Type: Contract / Full-Time
Job Summary
We are seeking an experienced AMS Lead to oversee Application Management Services (AMS) for mission-critical applications within the Telematics and Connected Car domain. The ideal candidate will have strong expertise in Site Reliability Engineering (SRE), Incident Management, Application Support, and Observability. This role requires leading incident response, driving service reliability, improving operational excellence, and collaborating with cross-functional teams to ensure high system availability and customer satisfaction.
Key Responsibilities
- Lead the Application Management Services (AMS) team for production support and operations.
- Manage the complete Incident Management lifecycle from Detection, Triage, Resolution, Recovery, and Root Cause Analysis (RCA).
- Act as the primary point of contact during critical production incidents and major outages.
- Coordinate with development, DevOps, infrastructure, cloud, and business teams to resolve production issues.
- Drive Service Reliability Engineering (SRE) best practices, including reliability, automation, and operational excellence.
- Monitor application health, availability, and performance using observability platforms.
- Perform trend analysis to identify recurring issues and implement permanent fixes.
- Lead post-incident reviews, prepare RCA reports, and drive corrective and preventive actions.
- Define and monitor SLAs, SLOs, KPIs, and operational metrics.
- Develop and improve operational runbooks, playbooks, and standard operating procedures.
- Identify opportunities for automation to reduce manual effort and improve system reliability.
- Support production releases, deployments, and change management activities.
- Mentor support engineers and ensure adherence to ITIL and operational best practices.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
- 8+ years of experience in Application Support, Production Support, or Application Management Services (AMS).
- 3+ years of experience leading Incident Management or Site Reliability Engineering (SRE) teams.
- Strong understanding of ITIL Incident, Problem, Change, and Service Management processes.
- Experience managing production incidents in high-availability environments.
- Excellent troubleshooting, analytical, and communication skills.
Required Technical Skills
- Site Reliability Engineering (SRE)
- Incident Management
- Major Incident Management
- Problem Management
- Root Cause Analysis (RCA)
- Application Management Services (AMS)
- Production Support
- Service Reliability
- High Availability Systems
- SLA / SLO / KPI Management
- Change Management
- Release Management
- Monitoring and Observability
- Dynatrace
- Grafana
- ELK Stack (Elasticsearch, Logstash, Kibana)
- Splunk
- Prometheus
- AppDynamics
- New Relic
- CloudWatch (AWS)
- Azure Monitor
- Log Analysis
- Performance Monitoring
- Alerting and Dashboarding
Domain Expertise
- Telematics
- Connected Car
- Automotive Cloud Platforms
- Vehicle Connectivity
- IoT Applications
- Fleet Management
- Vehicle Diagnostics
- Remote Vehicle Services
- Over-the-Air (OTA) Updates
- Infotainment Systems
- Automotive Data Platforms
Preferred Qualifications
- Experience with AWS, Azure, or Google Cloud Platform.
- Familiarity with Kubernetes, Docker, and containerized environments.
- Experience with CI/CD pipelines (Jenkins, GitLab CI, Azure DevOps).
- Scripting knowledge in Python, Shell, or PowerShell.
- Experience with REST APIs, microservices, and distributed systems.
- ITIL Foundation or SRE certification is highly preferred.
- Experience supporting automotive or connected mobility platforms.