3 days ago
Base Salary
$168k - $245k/yr
Responsibilities
- Operate and improve Kubernetes-based production infrastructure and deployment systems.
- Own installation, upgrades, troubleshooting, and lifecycle management for cloud and air-gapped customer deployments.
- Build and improve deployment observability, monitoring, logging, and alerting.
- Improve reliability, scalability, operational efficiency, and performance through automation and optimization.
- Participate in production incident response, root cause analysis, and reliability improvements.
- Tune infrastructure components, databases, and services for performance and resilience.
- Design and develop internal tooling using Python and/or Go.
- Debug production issues across Kubernetes, networking, storage, and application layers.
- Manage infrastructure using Terraform or similar Infrastructure as Code tools.
- Collaborate with software engineers and customers on secure, scalable, and reliable deployment architectures.
Requirements
- 7+ years of experience with a bachelor's degree, 4+ years with a master's degree, 1+ year with a PhD, or equivalent related experience.
- At least 4 years of experience in Site Reliability Engineering, Platform Engineering, Cloud Engineering, Infrastructure Engineering, or related fields.
- At least 3 years of experience operating Kubernetes in production and experience with Helm.
- Experience building and maintaining CI/CD platforms and deployment automation.
- Experience with AWS, GCP, or similar cloud platforms.
- Experience improving production reliability, scalability, and availability.
- Experience with monitoring, logging, observability, and alerting platforms.
- Strong scripting or programming skills in Python and/or Go.
- Experience with Infrastructure as Code tools such as Terraform.
- MLOps experience is preferred.
- Understanding of networking fundamentals including VPCs, DNS, routing, and load balancing.
- Experience operating and performance-tuning databases.
- Experience deploying and supporting cloud and air-gapped or on-premises environments.
- Strong debugging and troubleshooting skills across distributed systems.
Benefits
- Hybrid work requiring approximately two days per week onsite at Cisco offices in San Francisco, San Jose, or New York City.
- Medical, dental, and vision insurance, a 401(k) plan with Cisco matching contributions, paid parental leave, disability coverage, and basic life insurance, subject to eligibility.
- Potential eligibility for Cisco restricted stock unit grants.
- Paid holidays, birthday leave, year-end shutdown, personal wellness days, vacation or flexible vacation time, sick time, and family emergency leave, subject to applicable policies.
- Optional 10 paid volunteer days per calendar year.
- Eligibility for annual bonuses for non-sales roles, subject to Cisco policies.
Tech Stack
Categories
Site Reliability
About Cisco
Cisco designs and sells networking, security, and collaboration platforms for enterprises, service providers, and governments, spanning routers and switches, Wi‑Fi, firewalls, zero‑trust, observability, and cloud-managed IT (Meraki) plus Webex. Its business model mixes hardware, software subscriptions, and support/consulting services. Founded in 1984 and headquartered in San Jose, California, Cisco is a public company traded on Nasdaq and serves customers across data centers, campuses, and service provider networks.
