FIS

Lead Site Reliability Engineer

FIS
Apply
1 day ago
Bengaluru, IndiaStaff+
H1B Sponsor

Responsibilities

  • Own reliability outcomes for large-scale distributed payment and transaction-processing systems.
  • Define reliability architecture and standards for services, platforms, and infrastructure.
  • Design and evolve observability platforms covering metrics, logs, traces, SLOs, and SLIs.
  • Lead high-severity production incident response, root-cause analysis, and systemic remediation.
  • Drive SRE practices including error budgets, capacity modeling, resilience testing, graceful degradation, and operational readiness.
  • Architect automation and self-service platforms to reduce toil and operational risk and enable safe releases.
  • Partner with engineering, product, and platform leaders on architecture, cloud migration, disaster recovery, and platform evolution.
  • Mentor senior engineers and technical leads and improve organizational reliability maturity.

Requirements

  • Deep software engineering experience building and operating large-scale distributed, API-driven production systems.
  • Expertise in observability, alerting, and reliability engineering using tools such as Prometheus, Grafana, Datadog, Splunk, or ELK.
  • Strong knowledge of AWS, Azure, or GCP, infrastructure-as-code, platform automation, and cloud-native design patterns.
  • Significant experience operating mission-critical systems in Payments, FinTech, Banking, or similarly regulated environments.
  • Hands-on experience with Linux, RHEL, Windows, databases such as Oracle RDBMS, and complex enterprise systems.
  • Demonstrated leadership in incident management, post-incident reviews, and continuous reliability improvement.
  • Ability to operate at Staff-level scope, solve ambiguous problems, make trade-offs, and align multiple teams.
  • Strong automation and scripting skills with Python, Bash, Ansible, or similar tools are an advantage.
  • Experience building or scaling CI/CD platforms and release automation is an advantage.
  • Experience owning multi-team reliability or platform initiatives and modernizing legacy financial systems is an advantage.

Benefits

  • Full-time role working on high-impact payment platforms at massive scale.
  • Staff-level opportunity to define reliability strategy for critical financial systems.
  • Engineering culture emphasizing technical leadership, automation, excellence, and continuous learning.

Tech Stack

AnsibleAWSAzureBashDatadogGoogle Cloud PlatformGrafanaLinuxPrometheusPythonSplunkWindows

Categories

Site Reliability
FIS

About FIS

10,000+ employees

Unlocking financial technology. Bringing the world’s money into harmony. At FIS, we advance the way the world pays, banks, and invests. With decades of expertise, we provide financial technology solutions to financial institutions, businesses, and developers. Headquartered in Jacksonville, Florida, we’re a proud member of the Fortune 500® and the Standard & Poor’s 500® Index. Let's innovate together.

Contact me