Nebius

Senior ML Solutions Architect - Token Factory

Nebius
Apply
5 months ago
Remote, Singapore or Singapore, SingaporeSenior

Responsibilities

  • Design and implement LLM-based solutions using Nebius Token Factory inference services.
  • Build production-ready applications using serverless LLM APIs, including text, vision, audio, and domain-specific models.
  • Advise customers on prompt engineering, RAG architectures, model selection, and inference optimization.
  • Guide customers from proof of concept to production with attention to performance, reliability, and cost efficiency.
  • Collaborate with product and engineering teams to communicate customer feedback and shape the platform roadmap.

Requirements

  • At least 5 years of experience with ML/AI systems, including at least 2 years focused on LLMs and generative AI.
  • Deep knowledge of the LLM ecosystem, including model architectures and fine-tuning approaches.
  • Hands-on experience with prompt engineering, LLM pipeline development and evaluation, agentic frameworks, vector databases, RAG implementation, and LLM-powered applications using OpenAI, Anthropic, or open-source model APIs.
  • Strong Python programming skills.
  • Excellent communication skills for explaining technical concepts to diverse audiences.
  • Bonus experience includes vLLM, SGLang, TensorRT-LLM, Transformers, inference optimization, multimodal AI models, Docker, Kubernetes, and open-source ML/AI contributions.

Benefits

  • Remote work from Singapore is available.
  • Flexible working arrangements.
  • Comprehensive benefits package and professional growth opportunities.
  • Collaborative work environment focused on initiative and innovation.

Categories

AI ApplicationsSolutions Engineering
Nebius

About Nebius

1,001-5,000 employees

Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.

Contact me