10 hours ago
Sydney, AustraliaStaff+
Responsibilities
- Own the technical roadmap for ML infrastructure, including Kubernetes workflow orchestration, distributed training, batch inference, Ray-based real-time serving, and LLMOps.
- Lead and mentor a team of 3–4 mid-level and senior ML systems engineers and own technical hiring for the platform team.
- Design multi-cloud GPU capacity strategy, build foundational platform components, and develop the observability and detection stack.
- Write production Python and prototype high-risk platform capabilities before broader team adoption.
- Establish MLOps and AIOps practices spanning automation, infrastructure as code, CI/CD, and security across the AI lifecycle.
- Partner with Data Science and ML Engineering teams to convert platform needs into an actionable roadmap.
- Own service-level objectives and GPU cost efficiency across AWS and GCP, and lead major incident response.
Requirements
- 10+ years of software engineering experience, including at least 4 years building and operating large-scale infrastructure, platform engineering, or distributed systems in production.
- Experience leading technical projects and mentoring or managing engineering teams.
- Deep production expertise designing, building, and operating Kubernetes systems, including GPU scheduling.
- Expert-level Python and strong distributed-systems judgment around failure, backpressure, and cost at scale.
- Hands-on experience managing AWS and GCP infrastructure with infrastructure as code such as Terraform or Pulumi.
- Practical experience with Git, monitoring, alerting, and automated testing.
- Bachelor’s or Master’s degree in computer science, engineering, or a related technical field, or equivalent practical experience.
- Experience with workflow orchestration frameworks such as Argo Workflows, Kubeflow Pipelines, Flyte, Ray, or KubeRay is highly desirable.
- Experience designing large-scale GPU-intensive ML workloads and GPU capacity strategy across multiple clouds is highly desirable.
- Familiarity with model registries, feature stores, serving frameworks, vector databases, and RAG systems is highly desirable.
- Advanced Kubernetes capabilities, including custom operators and multi-cluster networking, are highly desirable.
- Experience with petabyte-scale data pipelines or geospatial and computer vision workloads is highly desirable.
- A track record of contributing to cloud-native or MLOps open-source projects is highly desirable.
Benefits
- Quarterly wellbeing day off and four additional annual “YOU” days.
- LinkedIn Learning, team hackathons, and pitch-fests.
- Wellbeing and technology allowance and annual flu vaccinations.
- Nearmap subscription and stocked kitchen.
- In-office lunch every Tuesday and Thursday at the Sydney CBD office, with showers available for cyclists and gym-goers.
- Hybrid flexibility for the role.
- Inclusive and supportive workplace culture.
Categories
About Nearmap
Nearmap provides high-resolution aerial imagery and AI-driven property intelligence delivered through web applications and APIs for insurers, government agencies, and AECO organizations. The company operates a recurring subscription model for frequently updated captures and analytics that support underwriting, claims, planning, and construction. Founded in 2009 and headquartered in Barangaroo, NSW, Nearmap was acquired by Thoma Bravo and serves customers across Australia, New Zealand, and North America.
