3 hours ago
Responsibilities
- Provision and manage cloud-native AI/ML infrastructure using Kubernetes, Docker, and GPU orchestration frameworks such as NVIDIA GPU Operator, Slurm, or Ray.
- Automate platform infrastructure with Terraform, Helm, and Ansible.
- Optimize GPU compute workloads, high-speed networking, and storage for model training and low-latency inference.
- Build and maintain pipelines for continuous model training, evaluation, packaging, and production deployment.
- Deploy large language models and generative AI workloads using Triton Inference Server, vLLM, or TensorRT-LLM.
- Automate model validation and monitor model drift, data drift, latency bottlenecks, and overall system performance.
- Monitor and optimize cloud spending across GPU and CPU clusters on AWS, GCP, or Azure.
- Implement auto-scaling, spot instance policies, and dynamic resource allocation to reduce infrastructure waste.
- Establish benchmarking and telemetry to measure unit economics and throughput for AI model training and serving.
- Implement end-to-end observability using Prometheus, Grafana, OpenTelemetry, Weights & Biases, or MLflow.
Requirements
- Hands-on production experience in DevOps, Site Reliability Engineering, or Platform Engineering, including experience with AI/ML infrastructure.
- Demonstrated experience deploying, scaling, and operationalizing machine learning models and LLMs in cloud-native production environments.
- Experience managing compute-intensive GPU infrastructure and high-performance computing environments.
- Advanced proficiency with Kubernetes, Docker, Helm, KubeFlow, and service meshes such as Istio.
- Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins.
- Experience with vLLM, Ray, MLflow, LangChain or LangSmith, DeepSpeed, or Hugging Face TGI.
- Experience with AWS, GCP, or Azure and Kubecost, including GPU cost optimization techniques.
- Strong skills in Python, Bash, or Go, with deep knowledge of Linux kernel tuning and performance monitoring.
Tech Stack
AnsibleAWSAzureBashDockerGitHub ActionsGoGoogle Cloud PlatformGrafanaHelmIstioJenkinsKubernetesLinuxMLflowPrometheusPythonTerraform
About Anaplan
Anaplan is a leading AI-driven scenario planning and analysis platform designed to optimize decision-making in today’s complex business environment so that enterprises can outpace their competition and the market. By building connections and collaboration across organizational silos, our platform intelligently surfaces key insights — so businesses can make the right decisions, right now. More than 2,500 global brands plan with Anaplan.
