
Lead Software Engineer
Wells Fargo18 hours ago
Bengaluru, IndiaStaff+
Responsibilities
- Own and improve production-system availability, performance, scalability, resilience, capacity planning, failover readiness, and disaster recovery.
- Define and manage SLIs, SLOs, error budgets, instrumentation standards, and reliability investments.
- Design and operate observability capabilities with Prometheus, Grafana, and Splunk, including dashboards, alerting, log aggregation, troubleshooting, and incident forensics.
- Develop Python automation, reusable SRE tooling, self-healing workflows, and CI/CD-integrated reliability processes.
- Apply AI/ML techniques to anomaly detection, log correlation, predictive capacity planning, noise reduction, and intelligent alerting.
- Serve as incident commander and senior escalation point for P1/P2 incidents, lead blameless post-incident reviews, and drive corrective actions.
- Partner with application, platform, cloud, and SRE teams on architecture, releases, migrations, modernization, and reliability-by-design.
- Ensure observability and automation meet security, audit, resilience, failover, and business-continuity requirements.
- Mentor SRE and software engineers and define standards for observability, automation, reliability, and incident response.
Requirements
- At least 5 years of software engineering experience, or equivalent demonstrated through work experience, training, military experience, or education.
- Strong Python proficiency for automation and tooling.
- Hands-on production experience with Grafana, Prometheus, and Splunk.
- Understanding of SLIs, SLOs, dashboards, alerting, observability, Linux, networking, and distributed systems.
- Experience with cloud platforms, Kubernetes or OpenShift, incident leadership, root-cause analysis, reliability initiatives, and large-scale or globally distributed systems.
- Experience building Prometheus exporters, advanced Grafana dashboards, Splunk searches, dashboards, alerts, and log pipelines.
- Experience operationalizing ML models for observability or AIOps and applying AI/ML concepts to monitoring, alerting, or operational analytics.
- Familiarity with Terraform, Ansible, CI/CD, and enterprise automation platforms.
- Ability to improve reliability and performance against SLOs, reduce alert noise, accelerate incident detection and recovery, and increase automation and self-healing adoption.
- Ability to lead complex technology initiatives, establish engineering standards, collaborate with technical experts, and mentor engineers.
Tech Stack
Categories
Site Reliability
About Wells Fargo
Wells Fargo & Company is a U.S.-based financial services firm providing consumer and commercial banking, mortgages, credit, and wealth/investment services to individuals, small businesses, and enterprises. It earns revenue from interest income and fees across retail banking, payments, lending, and capital markets. Founded in 1852 and headquartered in San Francisco, it is publicly traded on the NYSE (WFC) and operates nationally with offices in multiple countries.