
Staff/Sr. ML Infrastructure / Platform Engineer
Trend Micro24 hours ago
Taipei, TaiwanSenior / Staff+
Responsibilities
- Operate multi-model LLM serving infrastructure.
- Operate production Kubernetes clusters with NVIDIA GPU nodes.
- Manage GPU node lifecycles, including NVIDIA driver setup.
- Tune autoscaling policies to balance GPU cost and latency SLAs.
- Write and maintain Terraform/Terragrunt modules for AWS and GCP cloud environments.
- Package platform components and model deployments as Helm charts.
- Manage multi-environment configurations.
- Maintain Prometheus and Grafana monitoring.
- Build dashboards for GPU utilization, KV cache occupancy, TTFT/ITL latency, and cost per token.
- Set up alerting for SLA violations and out-of-memory events.
Requirements
- Experience operating production Kubernetes clusters with NVIDIA GPU nodes.
- Experience with LLM model serving infrastructure, GPU infrastructure, autoscaling, and inference optimization.
- Experience writing and maintaining Terraform/Terragrunt modules for AWS and GCP.
- Experience packaging platform components and model deployments as Helm charts.
- Experience maintaining monitoring and observability systems with Prometheus and Grafana.
- Bonus: experience with LoRA/PEFT fine-tuning workflows, MLflow, LoRA adapter CI/CD pipelines, SGLang, NVIDIA NIM, continuous batching, KV cache management, and speculative decoding.
Tech Stack
Categories
About Trend Micro
Trend Micro builds cybersecurity software and cloud services for enterprises, governments, and consumers, including endpoint, network, email, and cloud workload protection unified in the Trend Vision One platform. The company sells subscriptions and licenses with managed threat detection and response. Founded in 1988 and headquartered in Tokyo, it is a public company on the Tokyo Stock Exchange with global operations and longstanding partnerships with major cloud providers.