2 months ago
Hong Kong, Hong Kong +2 moreStaff+
Responsibilities
- Own the end-to-end ML platform strategy and define the roadmap for training pipelines, serving architecture, experiment management, and monitoring.
- Build foundational tooling that accelerates the path from ML experimentation to production.
- Solve distributed training, data residency, security, GPU efficiency, and sparse tensor implementation challenges.
- Design scalable serving systems for spiky production workloads across regions, customers, tasks, and verticals.
- Define interfaces between training, evaluation, and serving and remove bottlenecks for the ML team.
- Establish platform standards, documentation, and internal practices that raise the engineering discipline’s MLOps capability.
Requirements
- Significant technical experience running deep learning at scale and designing and operating ML systems used by other engineers.
- Hands-on expertise with training pipelines, distributed compute, model serving, model monitoring, and data quality frameworks.
- Experience building training data warehouses and preparing data systems for ML use.
- Experience setting ML platform standards, influencing engineering roadmaps, and aligning teams on complex infrastructure decisions.
- Strong proficiency in Python and PyTorch or an equivalent framework, with a passion for deep learning.
- Strong software engineering fundamentals in system design, code quality, and scalability.
- Experience with AWS, GCP, or Azure; Kubernetes and Docker; and custom on-premises or neocloud infrastructure.
- Experience with R&D prioritization and developing others through documentation and internal standards.
- Experience writing and optimizing custom CUDA kernels for deep learning training is desirable but not required.
Benefits
- Full relocation to Sydney, Australia.
- Visa sponsorship provided as part of the salary package.
- Competitive salary and meaningful ESOP.
- Fully flexible work environment with access to a stocked Redfern office.
- Regular office events.
- Opportunity to work on a complex, industry-leading product supporting climate resilience and critical infrastructure.
