17 days ago
Singapore, SingaporeSenior
Responsibilities
- Build and operate the ML infrastructure and platforms powering A1’s AI products.
- Design systems for model training, evaluation, deployment, inference, and experimentation.
- Build and optimize model-serving and inference infrastructure for high-throughput and low-latency workloads.
- Improve the reliability, scalability, latency, throughput, and cost efficiency of AI systems.
- Develop pipelines for data preparation, training, evaluation, model release, and continuous improvement.
- Build platforms and tooling that help AI engineers and researchers experiment, evaluate, and ship models faster.
- Develop evaluation and benchmarking infrastructure to measure model quality, performance, and regressions.
- Build production observability, monitoring, tracing, and alerting for AI/ML workloads.
- Identify bottlenecks across the ML stack and continuously improve system performance.
- Work with AI engineers, researchers, and product teams to deliver production-ready infrastructure.
Requirements
- Strong software engineering fundamentals and experience building production systems.
- Experience building ML infrastructure, platforms, or production machine learning systems.
- Experience with model deployment, inference, evaluation, or data pipelines.
- Strong understanding of distributed systems and system reliability.
- Ability to write clean, maintainable, production-quality code.
- Comfort working in ambiguous, fast-moving environments and taking ownership.
- Openness to experimentation and continuous improvement.
- Experience with Python, PyTorch or JAX, ML serving infrastructure, cloud infrastructure, ML/data pipelines, workflow orchestration, GPU infrastructure, performance tooling, vector databases, and retrieval infrastructure.
