7 days ago
Remote, United StatesStaff+
Base Salary
$114k - $165k/yr
Responsibilities
- Lead cross-functional reliability initiatives across multiple value streams and coordinate execution across teams.
- Define and evolve SRE best practices, tools, and methodologies across the organization.
- Architect enterprise-scale, multi-region AWS infrastructure balancing reliability, cost, performance, and security.
- Establish and operate SLOs, SLIs, and error budgets for critical services.
- Serve as incident commander for major incidents and lead postmortems with completed action items.
- Lead disaster recovery planning for critical financial services infrastructure.
- Build shared Terraform Infrastructure as Code foundations, including reusable modules, standards, and patterns.
- Design and implement production-scale Kubernetes patterns, including multi-tenancy, security policies, and advanced scheduling.
- Establish observability standards using Datadog and Splunk for metrics, logging, tracing, dashboards, and alerting.
- Set CI/CD standards and patterns, including pipeline-as-code and progressive delivery at scale.
- Lead chaos engineering, game days, and systematic reliability testing initiatives.
- Drive FinOps initiatives to optimize cloud spend while maintaining reliability targets.
- Lead a functional team of SREs without direct reports on projects and operational initiatives.
- Mentor SREs at multiple levels through coaching, design reviews, code reviews, and training sessions.
- Partner with Engineering, Product, and Security leadership on reliability priorities, zero-trust architecture, and compliance controls.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
- 7–10 years of Site Reliability Engineering experience or equivalent, with demonstrated technical leadership.
- Proven ability to lead technical teams and drive complex projects to completion.
- Expert AWS knowledge, including large-scale, multi-region architecture design.
- Deep Kubernetes expertise, including advanced features, security, and production-scale operations.
- Mastery of Terraform-based Infrastructure as Code, including shared platforms and frameworks.
- Production software engineering experience with Python and/or Go.
- Extensive experience with Datadog, Splunk, and monitoring at scale.
- Deep understanding of CI/CD principles and enterprise-grade pipeline implementation.
- Experience leading major incidents and conducting effective postmortems.
- Strong understanding of security, networking, and infrastructure design patterns.
- Experience mentoring engineers and building technical capabilities in teams.
- Preferred qualifications include prior Lead or Staff SRE/Operational Excellence leadership, financial services experience, compliance knowledge, AWS or Kubernetes certifications, experience implementing SRE in organizations with 500+ engineers, chaos engineering experience, open-source community leadership, service mesh experience, and technical speaking or writing experience.
Benefits
- Medical, dental, vision, and life insurance.
- 401(k) retirement plan with company matching contributions up to 6%, potential discretionary contributions, financial advisory services, and a broad investment lineup.
- Tuition reimbursement up to $5,250 per year.
- Generous paid time off upon hire, including paid company holidays and floating holidays.
- Paid volunteer time of 16 hours per calendar year.
- Paid parental leave, paid short- and long-term disability, and FMLA leave programs.
- Business Resource Groups open to all employees.
- Flexible work environment with reliable high-speed wired internet required for remote and hybrid positions; office work may be required if the home environment or connectivity is inadequate.
- Other necessary computer equipment will be provided.
Tech Stack
Categories
Site Reliability
