Mastercard is a global technology company that powers economies and empowers people through secure digital payments. They are seeking a Principal Site Reliability Engineer to ensure the reliability and scalability of their applications, supporting critical services for global operations while fostering a culture of operational excellence and proactive risk management.
Responsibilities:
- Serve as the primary contact responsible for the overall application health, performance, and capacity
- Support services before they go live through activities such as system design consulting, capacity planning and launch reviews
- Partner with the development and product team of a new application to establish the right monitoring and alerting strategy and create the framework to achieve zero downtime during deployment
- Performs operability and resilience design and implements and maintains highly reliable and scalable infrastructure
- Perform root cause analysis of incidents and collaborate with development teams to resolve issues
- Stay up to date with the latest technologies and trends in SRE and cloud computing
- Participate in on-call rotations and be available to respond to critical incidents
- Complete end-to-end run ownership of the product
- Practice sustainable incident response and blameless post-mortems while taking a holistic approach to problem solving and optimizing time to recover
- Automate data-driven alerts to proactively escalate issues. Work with development teams to establish SLOs and improve reliability
- Tackle complex development, automation, and business process problems
- Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead Mastercard in DevOps automation and best practices
- Performs operational and resilience Design and implements solutions for capacity planning and performance optimization
- Increase automation and tooling to reduce toil and manual intervention
- Analyses ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns
Requirements:
- BS degree in Computer Science or related technical field involving coding (e.g., physics or mathematics), or equivalent practical experience
- Strong understanding of DevOps principles, practices along with configuration management
- Experience in operational and resilience designing, building, and operating large-scale, distributed systems
- Appetite for change and pushing the boundaries of what can be done with automation. Be curious about new technology, infrastructure, and practices to scale our architecture and prepare for future growth
- Experience with algorithms, data structures, scripting, pipeline management, and software design
- Systematic problem-solving approach, analytical, coupled with strong communication skills and a sense of ownership and drive
- Interest in designing, analysing, and troubleshooting large-scale distributed systems
- Strong leadership and mentoring skills
- A passion for observability, automation and continuous improvement
- Willingness and ability to learn and take on challenging opportunities and to work as a member of matrix based diverse and geographically distributed project team
- Ability to balance doing things right with fixing things quickly. Flexible and pragmatic, while working towards improving the long-term health of the system
- Comfortable collaborating with cross-functional teams to ensure that expected system behaviour is understood and monitoring exists to detect anomalies