Cisco

Senior Kubernetes Platform Engineer - AI/ML Infrastructure

Cisco
Apply
5 days ago
Durham, NC, USASenior
H1B Sponsor

Base Salary

$139k - $204k/yr

Responsibilities

  • Architect, build, and operate large-scale on-premises Kubernetes platforms using OpenShift and Anthos.
  • Own Kubernetes control plane operations, etcd lifecycle management, cluster lifecycle management, and platform extensibility.
  • Design scalable, multi-tenant platforms for AI/ML and GPU-based workloads.
  • Enable and optimize machine learning training, inference, and LLM deployment pipelines on Kubernetes.
  • Build Kubernetes controllers, operators, CRDs, webhooks, and Go-based platform services.
  • Implement infrastructure as code and automation to improve platform scalability, consistency, and operational efficiency.
  • Drive AIOps capabilities including telemetry, anomaly detection, automated remediation, and self-healing systems.
  • Improve metrics, logs, traces, resource utilization, scheduling, and cluster performance.
  • Partner with ML engineers and data scientists to operationalize ML workflows.
  • Participate in on-call rotations, incident response, reliability work, and continuous operational improvement.
  • Mentor engineers and help define platform engineering standards and best practices.

Requirements

  • At least 8 years of software engineering experience.
  • At least 4 years of hands-on production Kubernetes experience with control plane ownership.
  • Strong experience operating on-premises or self-managed Kubernetes environments.
  • Deep expertise in etcd backup, restore, recovery, and upgrades.
  • Strong proficiency in Go and experience building Kubernetes controllers, operators, CRDs, and webhooks.
  • Deep understanding of Kubernetes internals, including the API server, scheduler, controller loops, and reconciliation.
  • Experience supporting AI/ML or GPU-based workloads on Kubernetes platforms.
  • Proven experience operating and debugging large-scale distributed systems.
  • Experience with on-call rotations and production incident management.
  • Preferred experience with bare-metal or enterprise on-premises infrastructure at scale.
  • Preferred exposure to Kubeflow, MLflow, distributed training systems, internal developer platforms, PaaS systems, AIOps, automated remediation, predictive operations, data-driven or ML-based reliability and capacity planning, and open-source contributions to Kubernetes or CNCF ecosystems.

Benefits

  • Hybrid role requiring some on-site work at the Research Triangle Park, North Carolina, Dallas, Texas, or Allen, Texas office.
  • Medical, dental, and vision insurance, 401(k) with Cisco matching contribution, paid parental leave, short- and long-term disability coverage, and basic life insurance.
  • Eligible employees may receive Cisco restricted stock units and annual bonuses subject to company policies.
  • Paid holidays, floating holiday, birthday day off, year-end shutdown, personal wellness days, vacation or flexible vacation time, sick time, and eligible family emergency leave.
  • Optional 10 paid volunteer days per year.
  • The application window is expected to close on September 10, 2026.

Tech Stack

KubernetesMLflowOpenShift
Cisco

About Cisco

10,000+ employees

Cisco is the worldwide technology leader that is revolutionizing the way organizations connect and protect in the AI era. For more than 40 years, Cisco has securely connected the world. With its industry leading AI-powered solutions and services, Cisco enables its customers, partners and communities to unlock innovation, enhance productivity and strengthen digital resilience. With purpose at its core, Cisco remains committed to creating a more connected and inclusive future for all.

Contact me