Cisco

Site Reliability Engineering Technical Leader

Cisco
Apply
1 day ago

Base Salary

$187k - $268k/yr

Responsibilities

  • Lead feasibility assessment, technical planning, and phased migration of eligible workloads from AWS-hosted Kubernetes clusters to the internal Kubernetes platform.
  • Operate and improve specialized Kubernetes workloads, including production readiness, incident response, performance improvements, and operational improvements.
  • Solve complex issues spanning applications, Kubernetes, Linux, networking, containers, and infrastructure.
  • Improve the scalability, reliability, security, performance, and operability of Kubernetes-hosted services.
  • Partner with Node Connectivity, firmware, cloud infrastructure, security, SRE, and product teams to coordinate cross-system changes.
  • Design, implement, and maintain production-quality Go software for distributed, concurrent, and networked systems.
  • Drive operational readiness through SLIs, SLOs, testing, on-call support, root-cause analysis, and long-term corrective actions.

Requirements

  • 10+ years of professional software, site reliability, or infrastructure engineering experience, including technical leadership of substantial production systems.
  • Experience designing, deploying, and operating large distributed services on Kubernetes.
  • 5+ years of programming experience in Go or a similar systems programming language.
  • Experience with production incident response, performance analysis, Kubernetes reliability practices, observability, automation, and operational readiness.
  • Knowledge of Linux and networking concepts including IPv4, IPv6, TCP, routing, DNS, and TLS.
  • Sound judgment in architecture, incident response, prioritization, and technical tradeoffs, with the ability to align stakeholders without formal authority.
  • Preferred experience operating Kubernetes across AWS, on-premises, or hybrid environments using AWS/EKS, Docker, Kustomize, GitLab CI/CD, or similar deployment systems.
  • Preferred experience with Kubernetes networking, ingress, load balancing, service discovery, traffic management, gRPC, Protocol Buffers, mutual TLS, PKI, VPNs, tunneling, or network security.
  • Preferred experience with OpenTelemetry, Prometheus, Datadog, distributed routing, packet processing, performance optimization, or failure testing.

Benefits

  • Starting salary range of $186,900.00 to $267,700.00 for U.S. and/or Canada locations, excluding incentive compensation, equity, and benefits.
  • Medical, dental, and vision insurance; 401(k) with Cisco matching contribution; paid parental leave; short- and long-term disability coverage; and basic life insurance, subject to eligibility.
  • Potential eligibility for Cisco restricted stock unit grants.
  • Paid holidays, birthday leave, year-end holiday shutdown, personal wellness days, vacation or flexible vacation time, sick time, family emergency leave, and optional paid volunteer days, subject to applicable policies.
  • Eligibility for annual bonuses for non-sales roles, subject to Cisco policies.

Tech Stack

AWSDatadogDockerGitLab CI/CDGogRPCKubernetesLinuxPrometheus

Categories

Site Reliability
Cisco

About Cisco

10,000+ employees

Cisco designs and sells networking, security, and collaboration platforms for enterprises, service providers, and governments, spanning routers and switches, Wi‑Fi, firewalls, zero‑trust, observability, and cloud-managed IT (Meraki) plus Webex. Its business model mixes hardware, software subscriptions, and support/consulting services. Founded in 1984 and headquartered in San Jose, California, Cisco is a public company traded on Nasdaq and serves customers across data centers, campuses, and service provider networks.

Contact me