6 days ago
Sydney, Australia or Melbourne, AustraliaSenior
Responsibilities
- Lead the design and implementation of large-scale, production-grade distributed systems for AI features.
- Own architecture decisions and direct distributed-systems strategy across the AI Products estate.
- Manage technical debt and operational reliability for production services and pipelines.
- Build interfaces and harnesses that transition machine-learning models from research into production.
- Design and build scalable distributed infrastructure for generative-AI features.
- Process web-scale data using Python, SQL, and distributed-processing engines.
- Deploy and operate production systems on AWS and Kubernetes.
- Collaborate with Applied Scientists to integrate reliable ML and LLM-based product features.
- Mentor junior engineers and improve the technical capability of the AI Products team.
Requirements
- At least 5 years of experience building and operating production Python or equivalent-language services at scale.
- Strong system-design and coding proficiency, with software engineering as the primary focus.
- Experience owning production services or pipelines, including operational or on-call responsibility, incident response, and technical-debt management.
- Deep understanding of distributed-processing principles and strong SQL capabilities.
- Experience integrating machine-learning models or large-language-model features into production systems.
- Familiarity with MLFlow, TensorFlow, or PyTorch is valued.
- Familiarity with Airflow or Prefect is valued.
- Prior experience applying or fine-tuning large language models in a product context is a desirable qualification.
Benefits
- Flexible hybrid working model combining office collaboration with remote work.
- Access to modern office spaces and team-aligned office and collaborative boost days.