
Senior Vice President, Site Reliability Engineer
BNY Mellon13 days ago
Responsibilities
- Design and implement observability across distributed systems using logs, metrics, and traces.
- Integrate and optimize AppDynamics, Dynatrace, Grafana, and Splunk for monitoring and telemetry.
- Develop dashboards, alerts, and telemetry frameworks that provide real-time system visibility.
- Automate repetitive operational work and build self-healing and auto-remediation solutions.
- Troubleshoot production issues, participate in incident triage and root cause analysis, and improve system stability.
- Define and measure service health through SLIs, SLOs, performance metrics, capacity planning, and architecture improvements.
Requirements
- At least 9 years of experience in Site Reliability Engineering or Software Engineering.
- Strong programming background in Java, preferably, or another modern programming language.
- Experience with at least one observability platform: AppDynamics, Dynatrace, Grafana, or Splunk.
- Hands-on experience supporting and troubleshooting production systems.
- Strong analytical and problem-solving skills with the ability to identify inefficiencies and drive automation.
- Preferred qualifications include distributed systems or microservices experience, CI/CD and DevOps familiarity, cloud platform or Kubernetes exposure, Python or Bash scripting, and knowledge of SRE concepts.
Tech Stack
Categories
Site Reliability