18 hours ago
Pune, India or Bengaluru, IndiaSenior
Responsibilities
- Design and develop reusable automation with PowerShell, Python, Ansible, APIs, and related technologies to reduce manual operational work.
- Implement infrastructure as code with Terraform and integrate automation into enterprise CI/CD and DevOps pipelines.
- Automate disaster recovery workflows including failover, failback, infrastructure validation, application validation, and customer validation.
- Improve monitoring, observability, alerting, health checks, remediation, and monitoring deployment using Grafana and other enterprise platforms.
- Support Azure, virtualized, and on-premises production infrastructure, including provisioning, configuration, patching, security hardening, and compliance validation.
- Provide hands-on Windows Server support, including troubleshooting services, Active Directory, Group Policy, DNS, networking, certificates, authentication, and operating system issues.
- Take technical ownership during Sev1/2/3 incidents, troubleshoot complex issues, communicate findings, perform root cause analysis, and implement permanent solutions.
- Identify reliability risks, operational patterns, and automation opportunities and drive improvements through completion.
- Embed AI into project delivery to improve productivity, automate routine activities, strengthen decision-making with trusted data, and support continuous process improvement.
Requirements
- At least 5 years of experience in Site Reliability Engineering, DevOps, cloud engineering, infrastructure engineering, or systems engineering.
- Strong hands-on scripting and automation experience with PowerShell, Python, Bash, or similar languages.
- Experience with CI/CD and enterprise DevOps platforms such as Azure DevOps, Jenkins, GitHub Actions, or GitLab CI.
- Experience with Ansible or similar configuration-management technologies and Terraform or other infrastructure-as-code technologies.
- Experience supporting Azure or comparable cloud infrastructure and business-critical production environments.
- Experience with monitoring and observability platforms such as Grafana and Prometheus.
- Understanding of high availability, disaster recovery, failover, and failback, plus working knowledge of Windows Server administration and troubleshooting.
- Strong production troubleshooting and incident-management skills with the ability to independently drive engineering solutions.
- Preferred experience includes Azure infrastructure, Azure DevOps, Azure Site Recovery, Kubernetes or AKS, ServiceNow or similar ITSM platforms, vulnerability-management and privileged-access technologies, and regulated financial-services environments.
- Familiarity with SRE practices including SLI/SLO concepts and reducing operational toil.
Benefits
- Unlimited vacation subject to local regulations and business priorities, with hybrid working arrangements and location-dependent paid leave policies.
- Employee Assistance Program, Wellbeing Champions, Gather Groups, and monthly well-being events.
- Medical, life, and disability insurance, retirement plans, lifestyle benefits, and other location-dependent benefits.
- Paid time off for volunteering and donation-matching opportunities.
- Access to inclusion communities and employee participation groups.
- Online learning and accredited courses through the Skills & Career Navigator tool.
- Participation in the Finastra Celebrates global recognition program and regular employee surveys.
Tech Stack
AnsibleAzureBashGitHub ActionsGitLab CI/CDGrafanaJenkinsKubernetesPowerShellPrometheusPythonTerraform
Categories
DevOpsSite Reliability
About Finastra
Finastra builds banking software for financial institutions, spanning core banking, lending, payments, and treasury, delivered on-premises and via cloud APIs. It sells licensed and SaaS products and an open platform (FusionFabric.cloud) used by over 7,000 customers, including 80% of the world’s top 50 banks, in more than 110 countries. Founded in 2017 and headquartered in Paddington, London, it is privately held.
