
Vice President, Site Reliability Engineering
BNY Mellon3 months ago
London, United KingdomStaff+
Responsibilities
- Design, develop, and deploy centralized engineering solutions that improve operational efficiency, reduce toil, and enhance resiliency across Production Services.
- Build full-stack applications and internal engineering tools, including backend services, APIs, automation layers, and user-facing interfaces.
- Develop self-service tooling, operational dashboards, alert enrichment, service recovery, incident-reduction, and workflow-automation solutions.
- Create reusable frameworks and components that standardize and accelerate operational processes across Production Services teams.
- Automate infrastructure, deployment, configuration, and runtime support activities using Ansible and Kubernetes.
- Define and continuously improve SLIs, SLOs, service health measures, monitoring, observability, and alerting capabilities.
- Apply AIOps capabilities to event correlation, anomaly detection, root cause analysis, predictive insights, and proactive issue prevention.
- Partner with engineering, infrastructure, production support, security, and risk teams to ensure solutions are secure, scalable, supportable, and aligned with enterprise standards.
- Identify manual or fragmented processes and convert them into efficient, automated, centrally consumable solutions.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical discipline, or equivalent practical experience.
- Strong full-stack development experience with hands-on Python and Java expertise for backend or service-layer engineering.
- Strong working knowledge of front-end development using React or Angular for operational or engineering interfaces.
- Experience designing and deploying end-to-end solutions from application development through production deployment and operational support.
- Experience in Site Reliability Engineering, Production Engineering, DevOps, Platform Engineering, or similar roles supporting business-critical applications.
- Strong foundation in Linux/Unix systems administration, scripting, troubleshooting, and infrastructure concepts.
- Hands-on experience with Ansible and Kubernetes in enterprise or production environments.
- Experience defining and operationalizing SLIs, SLOs, dashboards, alerts, and health indicators.
- Hands-on experience with Prometheus, Grafana, AppDynamics, and Splunk.
- Strong troubleshooting, analytical, problem-solving, verbal communication, and written communication skills.
- Preferred: experience building centralized internal platforms or shared engineering services for operational or enterprise users.
- Preferred: experience applying AIOps, machine learning, or intelligent automation in production support or reliability engineering.
- Preferred: exposure to CI/CD pipelines, infrastructure as code, API-driven automation, distributed systems, cloud-native platforms, or container-based architectures.
- Preferred: knowledge of Agile, DevOps, and SRE operating models, continuous improvement, and blameless post-incident practices.
- Preferred: ability to influence engineering standards and drive adoption of common tooling and automation patterns across teams.
Benefits
- The role is located in London.