5 days ago
Base Salary
$187k - $308k/yr
Responsibilities
- Define the technical roadmap for platform reliability, scalability, and operational excellence.
- Lead the architecture and evolution of deployment platforms for cloud and air-gapped customer environments.
- Establish SLOs, operational readiness, capacity planning, and resiliency review standards.
- Lead reliability and scalability initiatives across Kubernetes, deployment infrastructure, databases, and networking.
- Build automation, internal platforms, frameworks, and tooling to reduce operational toil.
- Lead production incident response, postmortems, root cause analysis, and long-term remediation.
- Partner with engineering leadership on platform architecture, deployment strategy, and production readiness.
- Mentor engineers and provide technical leadership through design reviews and operational guidelines.
- Collaborate with customers and internal teams on secure, scalable, and reliable cloud and on-prem deployment architectures.
Requirements
- 8+ years of experience with a bachelor's degree, 6+ years with a master's degree, or 3+ years with a PhD, or equivalent related experience.
- At least 6 years of experience in Site Reliability Engineering, Platform, Cloud, Infrastructure Engineering, or related fields.
- 5+ years of operating large-scale Kubernetes platforms in production.
- Experience designing highly available, scalable, and resilient distributed systems.
- Experience with AWS, GCP, or other public cloud platforms.
- Strong experience designing CI/CD platforms and deployment automation at scale.
- Expertise in observability, monitoring, alerting, capacity planning, and performance engineering.
- Strong programming skills in Python and/or Go.
- Deep experience with Infrastructure as Code using Terraform or similar tools.
- Experience with MLOps is preferred.
- Strong understanding of networking, distributed systems, storage, databases, and cloud architecture.
- Experience operating SaaS and enterprise/on-prem deployments.
- Demonstrated technical leadership across multiple engineering teams and experience leading incident management, postmortems, and long-term reliability initiatives.
Benefits
- Hybrid work arrangement requiring approximately two days per week onsite at Cisco offices in San Francisco, San Jose, or New York City.
- Medical, dental, and vision insurance, a 401(k) plan with Cisco matching contributions, paid parental leave, disability coverage, and basic life insurance, subject to eligibility.
- Potential eligibility for Cisco restricted stock unit grants.
- Paid holidays, floating holidays, birthday leave, year-end shutdown, personal wellness days, vacation, sick time, and family emergency leave subject to applicable policies.
- Optional 10 paid volunteer days per calendar year.
- Eligibility for annual bonuses for non-sales roles.
Tech Stack
Categories
DevOpsSite Reliability
About Cisco
Cisco designs and sells networking, security, and collaboration platforms for enterprises, service providers, and governments, spanning routers and switches, Wi‑Fi, firewalls, zero‑trust, observability, and cloud-managed IT (Meraki) plus Webex. Its business model mixes hardware, software subscriptions, and support/consulting services. Founded in 1984 and headquartered in San Jose, California, Cisco is a public company traded on Nasdaq and serves customers across data centers, campuses, and service provider networks.
