3 hours ago
Sydney, AustraliaSenior
Responsibilities
- Build and operate end-to-end ML training, evaluation, and deployment pipelines from data ingestion through customer environments.
- Automate model production workflows with CI/CD, artifact and model registries, reproducible environments, and infrastructure as code.
- Implement production monitoring, alerting, drift detection, data quality checks, and incident response for deployed models.
- Manage distributed training infrastructure, GPU clusters, scheduling, resource utilization, and troubleshooting across cloud, on-premises, and neocloud environments.
- Deploy and scale inference services across regions and customers while addressing latency, cost, and data residency requirements.
- Improve ML engineering tools and workflows and document standards as the team scales.
Requirements
- Hands-on experience building and operating production ML training pipelines, model serving, and monitoring systems.
- Strong Python skills and working knowledge of PyTorch or an equivalent framework.
- Practical experience with distributed training and GPU infrastructure, including scheduling, resource management, and diagnosing throughput or memory issues.
- Experience with AWS, GCP, or Azure; Kubernetes and Docker; and infrastructure as code.
- Experience with production model monitoring, data quality frameworks, and preparing training data for ML readiness.
- Sound engineering judgment, maintainable coding practices, and an understanding of failure modes and operational tradeoffs.
- Preferred experience includes CUDA or kernel-level optimization, point cloud or geospatial data, and deployments into regulated or air-gapped customer environments.
Benefits
- Competitive salary and meaningful ESOP.
- Fully flexible work environment with a stocked office in Redfern.
- Regular office events.
- Opportunity to work on a complex, innovative product supporting climate resilience and critical infrastructure.
- Equal employment opportunity and a diversity-focused workplace.
