
Senior Vice President, Site Reliability Automation Engineer
BNY Mellon4 months ago
Responsibilities
- Design and implement end-to-end observability across distributed systems using logs, metrics, and traces.
- Integrate and optimize AppDynamics, Dynatrace, Grafana, and Splunk.
- Develop dashboards, alerts, and telemetry frameworks and drive monitoring best practices.
- Automate repetitive operational work and build self-healing and auto-remediation solutions.
- Troubleshoot complex production issues and participate in incident management, triage, and root cause analysis.
- Define and measure service health using SLIs, SLOs, and performance metrics.
- Identify bottlenecks and reliability risks and contribute to performance optimization, capacity planning, and resilient system architecture.
Requirements
- 3–6 years of experience in Site Reliability Engineering or Software Engineering.
- Strong programming background in Java, preferred, or another modern programming language.
- Experience with at least one observability platform: AppDynamics, Dynatrace, Grafana, or Splunk.
- Hands-on experience supporting and troubleshooting production systems.
- Strong analytical and problem-solving skills and the ability to identify inefficiencies and drive automation.
- Preferred: experience with distributed systems or microservices architectures.
- Preferred: familiarity with CI/CD pipelines and DevOps practices.
- Preferred: exposure to cloud platforms and/or Kubernetes.
- Preferred: scripting experience with Python, Bash, or similar tools for automation.
- Preferred: knowledge of SRE concepts including observability, incident management, and reliability engineering.
Tech Stack
Categories
Site Reliability