
Lead Site Reliability Engineer, Vice President
Marks & Spencer Group plc2 hours ago
Base Salary
$150k - $190k/yr
Responsibilities
- Maintain application availability, latency, health, capacity, and operational efficiency across production and non-production environments.
- Improve the full service lifecycle through automation, deployment support, monitoring, capacity planning, and reliability engineering.
- Troubleshoot infrastructure, application, database, networking, hardware, and software issues, including log analysis, debugging, performance tuning, and root-cause analysis.
- Develop scripts, tools, utilities, and code changes to automate deployments and production-management tasks.
- Partner with application development, infrastructure, upstream data providers, and downstream consumers to improve support and reduce escalations.
- Own support documentation and knowledge bases while serving as a subject matter expert for supported applications.
- Identify risks, analyze environmental trends, stabilize systems, and lead reliability improvement efforts.
- Provide production support, manage requests and incidents, and participate in weekend on-call and 24/7 coverage.
Requirements
- 10+ years of experience in a production environment with a solid software development background, or 10+ years designing, developing, and implementing technical solutions or significant deep technical support experience.
- Strong experience with scripting languages such as Shell, Python, and Perl, plus the ability to diagnose problems, debug code, optimize performance, and automate routine tasks.
- Strong database skills with DB2, Sybase, or Oracle, and hands-on experience with Autosys or other batch-scheduling software.
- Experience with continuous integration and continuous deployment, on-demand environments using virtual machines and containers, and deployment automation.
- Knowledge of monitoring tools including Splunk, IP Soft, and Sockeye.
- Practical experience with Agile methodology such as Scrum.
- Knowledge of cloud-based deployment, security, and networking concepts in Azure and AWS.
- Understanding of database concepts, job schedulers, messaging, web services, UNIX/Linux/Windows operating systems, networking fundamentals, and application troubleshooting.
- Strong leadership, communication, analytical, organizational, and problem-solving skills, with the ability to work independently or across teams in high-pressure production environments.
Benefits
- Morgan Stanley offers comprehensive employee benefits and perks for employees and their families.
- The role includes weekend on-call rotation, offshore-time availability, and participation in 24/7 production support coverage.
- Employees have opportunities for career growth and movement across the business.
- The company supports diversity, inclusion, and employee development.
Categories
DevOpsSite Reliability