
Lead Engineer, Site Reliability Engineering
Mastercard20 hours ago
Singapore, SingaporeSenior / Staff+
Responsibilities
- Assess application infrastructure health, performance, monitoring, alerting, capacity, scalability, and resilience.
- Design and implement observability strategies, integrate infrastructure telemetry, and build dashboards for investigation and root-cause analysis.
- Lead incident reviews, identify root causes and compatibility risks, and drive remediation and mitigation strategies to closure.
- Use automation and AI technologies to improve proactive issue detection and enable self-healing capabilities.
- Develop testing and validation plans for environment builds, disaster recovery exercises, and post-maintenance activities.
- Partner with product and development teams on growth forecasting, architecture, SLIs/SLOs, and reliability throughout the service lifecycle.
- Lead training and knowledge-sharing initiatives across networking and infrastructure disciplines.
- Evaluate vendor hardware, firmware, and software upgrade roadmaps and conduct proof-of-concept testing.
- Participate in periodic on-call support for critical payment systems operating 24/7.
Requirements
- 5–10 years of experience in SRE or a related operations role, including 3+ years supporting e-commerce, financial services, or large-scale SaaS platforms.
- Strong infrastructure troubleshooting, analytical problem-solving, incident management, and root-cause analysis skills.
- Hands-on experience with observability and monitoring tools such as Splunk, Dynatrace, Prometheus, Grafana, ELK/EFK, and OpenTelemetry.
- Familiarity with SolarWinds and NetScout for network telemetry.
- Experience with packet-level debugging using tcpdump and Wireshark.
- Broad understanding of infrastructure supporting payment platforms, including platform services, networking, databases, and storage.
- Experience with Chef, Ansible, Terraform, JSON, and YAML.
- Ability to coordinate cross-functional troubleshooting and lead RCA processes to closure.
- Experience partnering with development teams on architecture, SLIs/SLOs, and embedding reliability into services.
- Ability to troubleshoot complex production issues and drive long-term corrective actions.
Benefits
- Periodic on-call responsibilities are required for critical payment systems operating year-round.
- The role supports Mastercard’s global payments network and critical national infrastructure.
- Employees must comply with Mastercard security policies, confidentiality requirements, breach-reporting obligations, and mandatory security training.
Tech Stack
Categories
Site Reliability
About Mastercard
Mastercard builds and operates a global payments network used by banks, merchants, fintechs, and governments, offering card processing, real-time payments, tokenization, and fraud/risk services. It generates revenue from transaction processing and assessment/service fees across more than 200 countries and territories. Founded in 1966 and headquartered in Purchase, New York, Mastercard is a public company listed on the NYSE (ticker: MA).