
Senior Principal Infrastructure Services (SRE Practice)
Northern Trust20 days ago
Pune, India or Bengaluru, IndiaStaff+
Responsibilities
- Lead the design and evolution of reliable, scalable, and performant distributed systems using SRE principles.
- Design and develop automation, tools, scripts, and platforms to reduce operational toil and human error.
- Lead production incident response, blameless post-incident reviews, root cause analysis, and long-term corrective actions.
- Architect and implement end-to-end observability using metrics, logs, traces, dashboards, alerts, SLIs, SLOs, and error budgets.
- Drive capacity planning, load testing, chaos engineering, fault injection, and continuous reliability improvement.
- Create system architecture documentation, runbooks, operational standards, and incident playbooks.
- Collaborate with product, development, platform, security, and operations teams to embed SRE practices into planning and delivery.
- Lead strategic reliability initiatives and mentor, coach, and develop high-performing technical teams.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience demonstrating advanced technical and leadership capabilities.
- 15+ years of progressive systems engineering experience emphasizing site reliability, large-scale systems operations, and software engineering in complex enterprise or cloud environments.
- 7+ years in a technical leadership role such as Team Lead or hands-on Technical Manager.
- Proficiency in one or more modern programming languages such as Python, Go, Java, or Ruby.
- Experience operating systems across hybrid environments, including on-premises infrastructure and public or private cloud platforms.
- Hands-on experience with containerization and container orchestration technologies.
- Experience designing and implementing observability solutions covering metrics, logs, traces, dashboards, and alerts.
- Deep understanding of distributed systems, networking fundamentals, failure modes, and modern software architectures.
- Experience designing and delivering Infrastructure as Code through automated CI/CD pipelines.
- Demonstrated success mentoring technical teams and leading cross-functional initiatives and complex projects.
- Hands-on expertise implementing automated remediation driven by observability signals and reliability metrics.
- Experience working in Agile and DevOps environments with strong communication, problem-solving, and stakeholder skills.
Benefits
- Flexible and collaborative work culture with opportunities for internal movement and access to senior leaders.
- Inclusive workplace with reasonable accommodation support.
- Opportunities to contribute to community philanthropy and volunteering.
- Role based in Northern Trust’s Pune office.