
Senior Lead Engineer, Platform Operations & Observability/Ingénieur principal, Opérations de plateforme et observabilité
McKessonBase Salary
$94k - $157k/yr
Responsibilities
- Lead monitoring and observability strategies across enterprise applications, platforms, and services.
- Design, implement, and optimize dashboards, alerts, telemetry, logging, and application performance monitoring solutions.
- Lead incident management, major incident response, escalation coordination, service restoration, and post-incident reviews.
- Conduct root cause analysis and lead corrective and preventive action planning.
- Partner with engineering teams to improve platform reliability, resilience, scalability, and operational readiness.
- Lead change management reviews and promote safe deployment and release practices.
- Execute regression and validation test plans after production deployments.
- Provide technical leadership, coaching, and mentorship to engineers.
- Influence architecture, automation, CI/CD, and operational excellence initiatives.
- Lead production readiness reviews and operational acceptance activities before major releases.
Requirements
- At least 7 years of professional experience in software engineering, site reliability engineering, platform engineering, DevOps, or related technical roles.
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or equivalent experience.
- Experience supporting large-scale production environments and enterprise applications.
- Hands-on experience with monitoring, observability, logging, alerting, and application performance monitoring tools.
- Proven experience leading incident management and production support activities.
- Experience performing root cause analysis and implementing preventive solutions.
- Experience with CI/CD, automation, DevOps practices, and software delivery pipelines.
- Experience with microservices, APIs, distributed systems, and cloud-based architectures.
- Experience with Dynatrace, Prometheus, Dotcom Monitor, or similar observability platforms is preferred.
- Experience with Kubernetes, containers, and cloud platforms such as Azure is preferred.
- Knowledge of ITIL-aligned incident, problem, and change management practices is preferred.
- Experience defining SLAs, MTTR, and service reliability metrics is preferred.
- Technical leadership and engineering-team mentorship experience is preferred.
- Experience in regulated or highly compliant environments and platform modernization is preferred.
- Strong analytical, troubleshooting, continuous improvement, and stakeholder communication skills.
- Understanding and mastery of AI tools such as Copilot and Rovo and AI agents is preferred.
Benefits
- Competitive compensation package with base pay of $94,400–$157,300, plus potential annual bonus or long-term incentive opportunities.
- Hybrid work arrangement with two mandatory days per week at the Dobrin Office, usually Monday and Wednesday.
- Participation in on-call support and out-of-hours production deployment rotations may be required.
- Equal employment opportunity and reasonable accommodation support are provided.
Tech Stack
Categories
About McKesson
Welcome to the official LinkedIn page for McKesson Corporation. We're an impact-driven healthcare organization dedicated to “Advancing Health Outcomes For All.” As a global healthcare company, we touch virtually every aspect of health. Our leaders empower our people to lead with a growth mindset and deliver excellence for our customers, partners, and the wellbeing of people, everywhere. We work with biopharma companies, care providers, pharmacies, manufacturers, governments, and others to deliver insights, products and services that make quality care more accessible and affordable. Delivering better health outcomes for our employees, our communities, and our environment. Every day, we strive to inspire and enable people to reach their full potential. To learn more about how #TeamMckesson helps improve care in every setting, visit: https://bit.ly/3xadvB0