
Site Reliability & Resilience -Senior manager
Ernst and Young3 months ago
Noida, IndiaStaff+
Responsibilities
- Define SLA, SLO, and SLI measures and agreements for products, platforms, and services.
- Design and implement resilient architecture and implementation practices throughout the product lifecycle.
- Design and implement observability solutions to track, report, and measure SLA adherence.
- Automate manual operational work and toil through automated systems and CI/CD improvements.
- Optimize IT infrastructure and operations costs through FinOps review, analysis, and improvement.
- Review and improve deployed products, product architecture, and inter-service dependencies, including simplification.
- Assess SRE maturity and provide transformation strategies and roadmaps for higher maturity.
- Implement SRE roadmaps and govern SRE solutions across enterprises and lines of business.
- Troubleshoot and tune performance, scalability, availability, application servers, JVMs, queries, and databases.
Requirements
- 15+ years of experience with software product engineering principles, processes, and systems.
- Hands-on experience with Java/J2EE, a web server such as Apache Tomcat or IBM HTTP Server, an application server such as Tomcat or WebSphere, and a major relational database such as Oracle.
- Experience with at least one CI/CD tool such as Azure DevOps, GitLab CI/CD, or Jenkins, and infrastructure-as-code tools such as Terraform, AWS CloudFormation, or Ansible.
- Experience with cloud technology such as AWS, Azure, or GCP; container and platform technologies such as Docker, Pivotal, Kubernetes, or OpenShift; and reliability tools such as Azure AppInsight, CloudWatch, or Azure Monitor.
- Experience with observability and APM tools such as Dynatrace or AppDynamics, metrics and log consolidation using Splunk, and the ELK Stack.
- Knowledge of queuing models, thread pools, request servicing processes, Linux/RHEL performance monitoring, web services, SOA, ESB/DataPower, and RESTful services.
- Knowledge of application design patterns, J2EE architectures, microservices, Spring Boot, cloud-native architectures, Java runtimes, garbage collection, and JVM parameter tuning.
- Experience generating and analyzing thread dumps and heap dumps, tuning queries, and understanding database architecture.
- Knowledge of at least one automation scripting language such as Python.
- Mastery of collaborative software development using Git, Jira, and Confluence.
- AI/ML and data analytics knowledge and experience are desirable.
Tech Stack
AnsibleAWSAzureDockerGitGitLab CI/CDGoogle Cloud PlatformJavaJenkinsKubernetesLinuxOpenShiftPythonSplunkSpring BootTerraform
Categories
DevOpsSite Reliability
About Ernst and Young
Ernst & Young (EY) provides audit/assurance, tax, consulting, strategy and transactions services to enterprises, financial institutions, and public‑sector clients. Structured as a global network of partner‑owned member firms, it sells professional services on a fee basis, including a dedicated Financial Services Organization for banking, insurance, and capital markets. Headquartered in London, EY was formed in 1989 from the merger of Ernst & Whinney and Arthur Young, and operates in 150+ countries.