6 months ago
Base Salary
$180k - $240k/yr
Responsibilities
- Scale distributed research workloads across multi-node, multi-GPU clusters using PyTorch Distributed and Ray Train, including VLA and world-model workloads.
- Optimize model and data parallelism, GPU utilization, low-level networking, and hardware performance across H100 and A100 clusters using NCCL, InfiniBand, and RoCE v2.
- Implement Kubernetes-native GPU scheduling and inference serving with NVIDIA GPU Operator, KubeFlow, TensorRT, ONNX Runtime, and Triton Inference Server.
- Build self-healing infrastructure and agent-driven automation for cluster health monitoring, infrastructure code review, resource optimization, and data curation.
- Design ML lifecycle infrastructure using MLFlow, Argo Workflows, and Kubernetes, including experiment tracking, feature-store integration, A/B testing, shadow deployments, and rollbacks.
- Develop infrastructure as code with Terraform and Helm and scale ETL and data pipelines using Apache Airflow, Kafka, Spark, S3, GCS, and Delta Lake.
- Establish observability for infrastructure and ML systems using Prometheus, Grafana, OpenTelemetry, and the ELK Stack, including latency, throughput, convergence, and drift metrics.
Requirements
- At least five years of experience in ML infrastructure, MLOps, or DevOps supporting high-scale compute environments.
- Deep understanding of multi-GPU training strategies including FSDP, DeepSpeed, and Ray Train, plus high-performance networking such as NCCL and InfiniBand.
- Expertise in Kubernetes, Terraform, and Helm, with a focus on GPU-native orchestration.
- Proven experience building or supporting agentic workflows for infrastructure or data automation using LLMs.
- Expertise with MLFlow, Argo Workflows, and Kubernetes, plus strong experience with Docker and Helm.
- Proficiency with Apache Airflow, Kafka, Spark, and GitOps automation.
- Proficiency in Python and Bash; Go or Rust experience is a plus.
- Familiarity with the Model Context Protocol, hybrid cloud and on-premises GPU cluster management, physical AI workloads, or LLM-based observability is a bonus.
Benefits
- Onsite five days per week at the Mountain View, California office.
- Inclusive workplace committed to diversity and equal opportunity.
Tech Stack
Apache AirflowApache KafkaApache SparkBashDockerGoGrafanaHelmKubernetesMLflowPrometheusPythonPyTorchRustTerraform
Categories
About Gatik AI
Gatik AI builds and operates autonomous middle‑mile trucking networks for retailers and consumer brands, delivered as a transportation‑as‑a‑service offering. Its Level 4, Class 3–7 trucks move B2B freight on short‑haul, repeatable routes, with commercial deployments in Texas, Arkansas, and Ontario, including a fully driverless service launched with Walmart in 2021. Founded in 2017 and headquartered in Santa Clara, it integrates software and hardware to run driverless deliveries between warehouses and stores.
