24 hours ago
London, United KingdomSenior
Responsibilities
- Design and build scalable machine learning platforms and infrastructure for model training, deployment, and serving.
- Develop highly available backend services supporting recommendation, search, and personalisation experiences.
- Build and maintain CI/CD pipelines and automated deployment processes for machine learning and data products.
- Design batch and real-time inference architectures using cloud-native technologies.
- Improve reliability, resilience, performance, monitoring, observability, and automation across ML workloads.
- Build tooling and frameworks that help data scientists and ML engineers deploy models safely and efficiently.
- Own production services and infrastructure, including incident management and root cause analysis.
- Optimise distributed compute workloads and resource utilisation across cloud environments.
- Drive Infrastructure-as-Code adoption and platform standardisation across ML systems.
- Contribute to architecture decisions for recommendation, search, and AI platforms.
- Mentor engineers, promote software engineering best practices, and influence the long-term ML platform strategy.
Requirements
- Strong software engineering fundamentals and experience designing and building production systems at scale.
- Experience developing distributed systems, microservices, or high-throughput backend platforms.
- Strong programming skills in Python, Java, Kotlin, Go, Scala, or similar languages.
- Experience operating services in AWS, Azure, or GCP environments.
- Hands-on experience with Kubernetes, containerisation, and cloud-native technologies.
- Experience implementing CI/CD pipelines, automated deployments, monitoring, alerting, and observability.
- Knowledge of Infrastructure-as-Code tools such as Terraform, Pulumi, or CloudFormation.
- Experience building reliable, resilient, scalable, and performance-focused systems.
- Experience with data-intensive systems, streaming technologies, or large-scale distributed processing platforms.
- Exposure to machine learning systems, model serving, feature stores, training infrastructure, or MLOps practices.
- Experience supporting large-scale model training and inference workloads.
- Knowledge of vector search, ranking systems, retrieval architectures, recommendation platforms, or customer-facing data products is advantageous.
- Exposure to LLMs, generative AI, production AI systems, internal developer platforms, or engineering enablement tooling is advantageous.
- Comfort providing technical leadership, mentoring engineers, influencing architecture, and collaborating in cross-functional product teams.
Benefits
- Employee discount and access to employee sample sales.
- 25 days of paid annual leave plus an additional celebration day.
- Discretionary bonus scheme.
- Private medical care scheme.
- Flexible benefits allowance that can be taken as cash or used for other benefits.
- Personalised learning and in-the-moment development experiences.
Categories
About ASOS
ASOS is a global online fashion and beauty retailer for 20‑somethings, selling own‑brand and third‑party labels via its website and mobile apps. The company earns revenue from direct‑to‑consumer e‑commerce, shipping to customers worldwide and offering marketplace services for brands. Founded in 2000 and headquartered in London, ASOS is a public company listed on the London Stock Exchange and runs in‑house content and fulfillment operations at scale.
