3 months ago
Sydney, AustraliaStaff+
Responsibilities
- Own the ML platform strategy and multi-year roadmap for training pipelines, serving architecture, experiment management, and monitoring.
- Build foundational tooling and infrastructure that accelerates the path from ML experimentation to production.
- Solve distributed systems challenges involving distributed training, data residency, security requirements, GPU performance, sparse tensors, and architecture bottlenecks.
- Design flexible serving systems that support spiky production workloads across regions, customers, tasks, and verticals.
- Define interfaces between training, evaluation, and serving and remove delivery bottlenecks for the ML team.
- Set ML platform standards, influence engineering roadmaps, and lead, mentor, and coach ML engineers.
Requirements
- Significant hands-on experience building and operating training pipelines, distributed compute, model serving, and monitoring systems at scale.
- Strong proficiency writing and optimizing custom CUDA kernels for deep learning training, ideally involving non-text and non-image data, with experience diagnosing sparse architecture bottlenecks.
- Experience with production model monitoring, data quality frameworks, and preparing training data warehouses for ML use.
- Ability to set ML platform standards and align teams on complex infrastructure decisions without direct authority.
- Proficiency in Python and PyTorch or an equivalent framework, with strong system design skills and an R&D foundation.
- Experience with AWS, GCP, or Azure; Kubernetes and Docker; and on-premises or neocloud environments.
- Track record of leading, mentoring, and coaching ML engineers through hands-on guidance, documentation, and internal standards.
Benefits
- Competitive salary and meaningful ESOP.
- Fully flexible work environment with a stocked office in Redfern.
- Regular office events.
- Opportunity to work on a complex, innovative product using ML for critical infrastructure and climate resilience.
