Nvidia

Senior DevOps Engineer, AIOps

Nvidia
Apply
2 days ago
Tel Aviv-Yafo, IsraelSenior

Responsibilities

  • Own the DevOps, infrastructure, security, release, and reliability lifecycle from development environments and CI/CD through deployment and sustained operations.
  • Build and operate Kubernetes environments and Helm deployments for Python, FastAPI, Node.js, and React microservices across SaaS and on-premises environments.
  • Develop GitLab CI/CD pipelines with automated testing, container builds, vulnerability scanning, and versioned image and Helm chart publication through JFrog Artifactory.
  • Automate infrastructure provisioning, configuration, upgrades, and operational workflows using infrastructure-as-code and configuration tooling.
  • Operate PostgreSQL, Temporal workflow services, and S3-compatible object storage, including capacity planning, backups, recovery testing, and migrations.
  • Improve release reliability through deployment validation, reduced-downtime strategies, persistent-state protection, and recovery planning.
  • Implement observability and security using OpenTelemetry, Datadog/Grafana, Langfuse, secrets management, identity integration, TLS, Kubernetes RBAC, network policies, and container hardening.
  • Partner with software and AI engineers to troubleshoot distributed systems, investigate incidents, define reliability targets, and improve platform performance and resource efficiency.

Requirements

  • Bachelor’s degree in Computer Science, Software Engineering, or a related field, or equivalent experience.
  • At least 5 years of experience in DevOps, site reliability engineering, or platform engineering supporting distributed applications and microservices.
  • Hands-on experience with Kubernetes, Docker, and Helm, including networking, storage, workload scheduling, scaling, and troubleshooting.
  • Strong Linux administration skills and proficiency in Python and Bash for automation.
  • Experience with infrastructure as code and configuration tooling such as Terraform and Ansible.
  • Experience building and maintaining CI/CD pipelines, container registries, artifact management, automated quality gates, and secure release practices.
  • Experience operating PostgreSQL or comparable relational databases, including SQL, migrations, backup and restore, and performance troubleshooting.
  • Strong networking and observability fundamentals covering TCP/IP, DNS, HTTP, TLS, load balancing, ingress, metrics, logs, traces, dashboards, and alerting.
  • Understanding of secure infrastructure operations and incident response, with demonstrated ownership, collaboration, and prioritization.
  • Preferred experience includes operating AI applications, agent platforms, or LLM services; Temporal, LangGraph, Model Context Protocol, Langfuse, ClickHouse, Redis/Valkey, or S3-compatible storage; OpenTelemetry, Datadog APM, or Prometheus/Grafana; self-hosted Kubernetes, OpenShift, Kubernetes operators, CloudNativePG, GPU clusters, or AI data centers; and AMD64/ARM64 container images, BuildKit pipelines, and software supply-chain security.

Benefits

  • Competitive salaries and a generous benefits package.
  • Hybrid work arrangement, as indicated by the #LI-Hybrid designation.

Tech Stack

AnsibleBashBuildkiteClickHouseDatadogDockerFastAPIGitLab CI/CDGrafanaHelmKubernetesLinuxNode.jsOpenShiftPostgreSQLPrometheusPythonReactRedisSQLTerraform

Categories

Nvidia

About Nvidia

10,000+ employees

Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.

Contact me