Wells Fargo

Lead Platform Reliability Engineer

Wells Fargo
Apply
1 day ago
Minneapolis, MN, USA +2 moreStaff+

Base Salary

$119k - $224k/yr

Responsibilities

  • Serve as the reliability engineering expert for a primary Network, Middleware, Database, or Storage domain while partnering across adjacent infrastructure disciplines.
  • Lead investigation and resolution of complex production incidents, identify root causes, and implement long-term corrective actions.
  • Apply SRE practices including SLIs, SLOs, error budgets, incident analysis, and reliability measurement.
  • Lead capacity analysis, forecasting, utilization reviews, performance analysis, and scaling-risk mitigation.
  • Drive reliability improvements through observability, automation, performance optimization, resiliency engineering, and platform hygiene.
  • Design automation, operational tooling, API integrations, self-healing capabilities, and automated remediation to reduce operational toil.
  • Define and enhance observability standards for metrics, logging, tracing, alerting, and service health monitoring.
  • Partner with engineering, infrastructure, application, cloud, and operations teams to improve platform performance and availability.
  • Lead blameless post-incident reviews and turn recurring operational issues into measurable engineering improvements.
  • Identify reliability risks, communicate recommendations to technical leaders and senior stakeholders, and mentor engineers on reliability and operational excellence.

Requirements

  • At least 5 years of Systems Engineering, Infrastructure Engineering, Platform Engineering, Technology Architecture, or equivalent experience demonstrated through work experience, military experience, training, or education.
  • At least 5 years supporting and engineering enterprise-scale production environments.
  • At least 5 years of hands-on expertise in Network, Middleware, Database, or Storage Engineering.
  • Experience applying SRE principles, supporting highly available mission-critical environments, and troubleshooting complex issues across multiple technology domains.
  • Experience with capacity planning, resiliency engineering, fault tolerance, disaster recovery, performance optimization, and observability or monitoring platforms.
  • Experience building dashboards, alerts, service health indicators, operational reporting, and automated remediation solutions.
  • Strong automation and scripting experience using Python, Bash, PowerShell, or similar technologies.
  • Familiarity with Git-based development practices, CI/CD pipelines, infrastructure automation, and Infrastructure as Code tools such as Ansible or Terraform.
  • Experience leading major incident reviews, influencing technical direction, mentoring engineers, and communicating technical concepts as business-focused outcomes.

Benefits

  • Hybrid work schedule.
  • Health benefits, 401(k) plan, paid time off, disability benefits, life insurance, critical illness insurance, and accident insurance.
  • Parental leave, critical caregiving leave, commuter benefits, tuition reimbursement, dependent-child scholarships, adoption reimbursement, and discounts and savings.
  • Posting end date is September 11, 2026, though the posting may come down early due to applicant volume.

Tech Stack

AnsibleApache KafkaBashGitGrafanaMicrosoft SQL ServerMongoDBPostgreSQLPowerShellPrometheusPythonSplunkTerraform

Categories

DevOpsSite Reliability
Wells Fargo

About Wells Fargo

10,000+ employees

Wells Fargo & Company is a U.S.-based financial services firm providing consumer and commercial banking, mortgages, credit, and wealth/investment services to individuals, small businesses, and enterprises. It earns revenue from interest income and fees across retail banking, payments, lending, and capital markets. Founded in 1852 and headquartered in San Francisco, it is publicly traded on the NYSE (WFC) and operates nationally with offices in multiple countries.

Contact me