24 hours ago
Pune, IndiaSenior
Responsibilities
- Design and maintain monitoring and observability solutions for infrastructure, applications, and customer experience.
- Build automation frameworks and Infrastructure as Code solutions for efficient, consistent cloud deployments.
- Ensure high availability, reliability, scalability, and performance of critical production systems.
- Lead incident triage, root cause analysis, recovery, and post-incident reviews.
- Perform capacity planning, performance tuning, and infrastructure optimization.
- Maintain and optimize CI/CD pipelines for reliable and secure software delivery.
- Implement platform security controls and compliance best practices with security teams.
- Develop and improve disaster recovery, backup, and business continuity strategies.
- Collaborate with engineering, DevOps, QA, and product teams to meet service-level objectives.
- Participate in on-call rotations and support critical production environments.
Requirements
- Require 5+ years of experience in Site Reliability Engineering, Production Support, Platform Engineering, DevOps, Cloud Operations, or a related field.
- Require hands-on experience with AWS, Microsoft Azure, or Google Cloud Platform.
- Require knowledge of Infrastructure as Code and automation tools such as Terraform and Ansible.
- Require experience supporting web applications, APIs, distributed systems, and modern software architectures.
- Require proficiency with monitoring and observability tools such as Prometheus, Grafana, or Datadog.
- Require experience with logging and analytics solutions such as Splunk or ELK Stack.
- Require scripting and automation skills using Python, Bash, or similar languages.
- Require experience designing and optimizing CI/CD pipelines using Jenkins, GitLab CI/CD, Azure DevOps, or related tools.
- Require knowledge of Docker and container orchestration platforms.
- Require experience with incident management, root cause analysis, and enterprise production support.
- Preferred qualifications include large-scale SRE practices, Kubernetes, cloud-native and microservices architectures, operational readiness assessments, post-mortems, disaster recovery, high-availability design, resiliency engineering, and relevant industry certifications.
Benefits
- Full-time role located in Pune.
- Collaborative, flexible, and respectful work environment.
- Competitive salary and benefits supporting lifestyle and wellbeing.
- Varied and challenging work supporting technical growth.
Tech Stack
AnsibleAWSAzureBashDatadogDockerGitLab CI/CDGoogle Cloud PlatformGrafanaJenkinsKubernetesPrometheusPythonSplunkTerraform
Categories
Site Reliability
About FIS
FIS builds financial technology for banks, merchants, and capital-markets firms, including core banking platforms, payment processing, fraud/risk tools, and investment/treasury systems. It sells enterprise software, cloud services, and outsourced processing via long-term contracts and transaction-based fees. Founded in 1968 and headquartered in Jacksonville, Florida, FIS is a Fortune 500 public company trading on the NYSE (FIS) with customers worldwide.
