13 hours ago
Singapore, SingaporeSenior
Responsibilities
- Lead customer transitions from proof of concept to production and drive time-to-production and time-to-value.
- Ensure customer workloads are deployed, stable, scalable, performant, reliable, and cost-efficient.
- Work directly with customer engineering teams to understand architectures, use cases, and technical challenges.
- Monitor and improve latency, throughput, cost efficiency, and reliability while proactively resolving bottlenecks.
- Act as the primary technical contact for production issues and coordinate incident resolution with internal teams.
- Provide best-practice guidance, manage multiple customers and priorities, and communicate clearly during high-pressure situations.
- Identify opportunities to optimize and expand usage and provide structured feedback to Product and Infrastructure teams.
Requirements
- Practical knowledge of inference frameworks such as vLLM, TensorRT, or similar.
- Solid understanding of cloud or infrastructure systems, distributed systems or high-load applications, and AI/ML workloads including LLM inference.
- Ability to troubleshoot and reason about system performance.
- Experience working directly with technical customers such as engineers or ML teams.
- Ability to communicate complex technical topics clearly and effectively.
- Strong ownership, proactive execution, structured problem-solving, and ability to manage multiple customers and priorities.
- Preferred experience includes GPU workloads or AI infrastructure and a background in solutions engineering, site reliability engineering, or technical support in B2B environments.
Benefits
- Competitive compensation, career growth, and learning opportunities.
- Flexibility, ownership, a collaborative and innovative culture, and an international environment.
- Opportunity to work on impactful AI projects.
- Remote work is available from Singapore.
Categories
Forward Deployed
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
