6 months ago
Base Salary
$180k - $240k/yr
Responsibilities
- Architect and maintain mission-critical Kubernetes clusters optimized for GPU and TPU workloads.
- Implement Kubernetes-native GPU scheduling with NVIDIA GPU Operator and optimize hardware utilization.
- Manage infrastructure as code using Terraform, Helm, and cloud-native tools.
- Deploy automated infrastructure monitoring and triage for hardware failures and NCCL timeouts using AI agents.
- Build sensor-data pipelines with Apache Airflow, Kafka, and Spark.
- Implement GitOps and CI/CD workflows with ArgoCD and GitLab CI/CD for infrastructure and model artifacts.
- Maintain infrastructure and model-serving observability with Prometheus, Grafana, and OpenTelemetry.
- Develop agent-driven workflows, including automated Terraform pull-request review and Kubernetes resource-limit recommendations.
- Integrate MLFlow and feature stores for experiment and model tracking.
- Automate model lifecycles from training through simulation using Airflow and Kubernetes.
- Support model deployment with Triton Inference Server, Ray Serve, and ONNX Runtime.
- Enable multi-node distributed training with PyTorch Distributed, Ray Train, and Horovod.
- Optimize NCCL, InfiniBand, and RoCE v2 communication for large-scale training workloads.
- Tune performance across multi-node GPU clusters for FSDP and DeepSpeed workloads.
Requirements
- 5+ years of experience in cloud infrastructure, DevOps, or MLOps supporting high-scale compute environments.
- Deep expertise in Kubernetes, Helm, and container orchestration.
- Strong experience with Apache Airflow, Argo Workflows, MLFlow, and Terraform.
- Practical experience supporting distributed systems and frameworks such as Ray and PyTorch Distributed.
- Proficiency in Python and Bash scripting, with a solid understanding of IAM and RBAC.
- Bonus: deep understanding of FSDP and DeepSpeed.
- Bonus: experience building agentic workflows with LangGraph or AutoGen for infrastructure automation or data curation.
- Bonus: familiarity with Model Context Protocol (MCP).
Benefits
- Onsite five days per week at the Mountain View, California office.
- The role supports autonomous transportation technology and AI infrastructure.
- The company emphasizes collaboration, inclusion, professional growth, and sustainability.
Tech Stack
Apache AirflowApache KafkaApache SparkBashGitLab CI/CDGrafanaHelmKubernetesMLflowPrometheusPythonPyTorchTerraform
Categories
About Gatik AI
Gatik AI builds and operates autonomous middle‑mile trucking networks for retailers and consumer brands, delivered as a transportation‑as‑a‑service offering. Its Level 4, Class 3–7 trucks move B2B freight on short‑haul, repeatable routes, with commercial deployments in Texas, Arkansas, and Ontario, including a fully driverless service launched with Walmart in 2021. Founded in 2017 and headquartered in Santa Clara, it integrates software and hardware to run driverless deliveries between warehouses and stores.
