AMD

Devops Platform Engineer

AMD
Apply
2 days ago
San Jose, CA, USASenior

Responsibilities

  • Design, build, and operate scalable Kubernetes platform capabilities across development, test, and production environments.
  • Create reusable infrastructure-as-code modules, patterns, and automated workflows.
  • Develop platform observability for metrics, logs, traces, dashboards, alerting, SLOs, and runbooks.
  • Improve reliability, capacity management, security posture, and cost efficiency through automation.
  • Build infrastructure for AI/ML workloads, including model serving, GPU-enabled compute, data access, and workload isolation.
  • Enable developer self-service through platform APIs, templates, documentation, golden paths, and CI/CD integrations.
  • Partner with application, security, infrastructure, and AI teams on standards and delivery improvements.
  • Investigate production issues, lead root-cause analysis, and implement preventative improvements.
  • Contribute to platform roadmaps, architecture reviews, and engineering standards.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience.
  • At least 6 years of experience in software engineering, DevOps, SRE, cloud infrastructure, or platform engineering.
  • Strong hands-on experience operating and automating Kubernetes environments.
  • Experience with container technologies and deployment tools such as Docker, Helm, Kustomize, Argo CD, or Flux.
  • Experience implementing infrastructure as code with Terraform, Pulumi, CloudFormation, Ansible, or comparable tools.
  • Practical experience with observability platforms and metrics, logging, distributed tracing, alerting, dashboards, and SLOs.
  • Proficiency in at least one of Python, Go, Bash, or TypeScript.
  • Experience with CI/CD systems and automated software delivery.
  • Strong troubleshooting skills across Linux, networking, containers, cloud or on-premise infrastructure, and distributed systems.
  • Clear written and verbal communication skills and the ability to collaborate across engineering disciplines.
  • Preferred experience includes internal developer platforms, Prometheus, Grafana, OpenTelemetry, Loki, Tempo, Elastic, Datadog, Splunk, IBM LSF, NetApp Storage, GitOps, AI/ML infrastructure, model serving, GPU scheduling, Kubernetes operators, service mesh, API gateways, ingress, multi-cluster Kubernetes, platform security, and semiconductor or EDA environments.

Tech Stack

AnsibleArgo CDBashDatadogDockerGoGrafanaHelmKubernetesLinuxPrometheusPythonSplunkTerraformTypeScript

Categories

AMD

About AMD

10,000+ employees

AMD designs and sells CPUs, GPUs, and adaptive/embedded computing products for PCs, data centers, gaming, and edge devices. Its portfolio includes Ryzen and EPYC processors, Radeon and Instinct graphics, and adaptive SoCs from its Xilinx acquisition, sold to OEMs, cloud providers, and device makers. Founded in 1969 and headquartered in Santa Clara, it is a public company on NASDAQ and supplies semi-custom chips for major game consoles.

Contact me