9 hours ago
Melbourne, AustraliaSenior
Responsibilities
- Build and own model-serving infrastructure, including deployment pipelines, request routing, autoscaling, fallback paths, and controlled rollouts and rollbacks.
- Manage GPU clusters and scheduling policies across inference, training, and evaluation workloads.
- Profile production workloads and improve inference latency, throughput, memory efficiency, and cost.
- Support distributed training and model iteration through reliable job launching, artifact management, checkpointing, and model promotion to serving.
- Build dashboards, alerts, and tracing for model latency, queueing, errors, GPU health, memory pressure, and workload performance.
- Own production reliability by defining service objectives, investigating incidents, and developing recovery procedures, automated checks, and runbooks.
- Track GPU usage, idle capacity, capacity forecasts, and inference cost by model and workload.
- Automate provisioning, configuration, benchmarking, and releases so the model team can deploy and evaluate new models independently.
Requirements
- At least one year of experience building and operating infrastructure for large language models, including model deployment, inference serving, or distributed training.
- Experience deploying and maintaining LLMs or demanding machine-learning workloads with systems such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or comparable tools.
- Practical experience managing GPU workloads with Kubernetes, Slurm, or an equivalent platform, including scheduling, resource allocation, capacity planning, and failure recovery.
- Ability to use traces, metrics, and profiling tools to diagnose compute, memory, communication, and scheduling bottlenecks.
- Proficiency in Python and comfort with backend or systems development in Go, C++, Rust, or a comparable language.
- Practical experience with Linux, containers, deployment automation, distributed services, production incidents, monitoring, and recoverable releases.
- Preferred experience with PyTorch FSDP, Megatron, reinforcement-learning infrastructure, inference-engine tuning, MoE models, quantization, speculative decoding, or prefill/decode disaggregation.
- Preferred familiarity with GPU interconnects, NCCL, RDMA, topology-aware scheduling, multi-node communication, CUDA or Triton kernel development, cluster operators, scheduling integrations, and multi-region or multi-provider infrastructure.
Benefits
- A $1,000 annual learning and development budget.
- A $150 per month health and wellness allowance.
- A $500 home office budget.
- 26 weeks of paid primary parental leave and 18 weeks of paid secondary parental leave.
- Fertility support up to $10,000.
- Four weeks of work from anywhere per year and equity.
- Flexible scheduling focused on outcomes, with emphasis on sustainable performance and mental health.
Categories
About Heidi
Heidi builds an AI care partner for clinicians that automates documentation and workflows, including an AI medical scribe, Evidence for point‑of‑care research, and Comms for patient coordination. It sells a freemium and enterprise SaaS platform to healthcare providers and health systems. Headquartered in Melbourne and privately held, Heidi reports supporting over 2.7 million patient interactions each week in 110 languages across 190 countries.
