6 months ago
Remote, EMEASenior
Responsibilities
- Design and implement LLM-based solutions using Nebius Token Factory inference services.
- Build production-ready applications with serverless LLM APIs, including text, vision, audio, and domain-specific models.
- Provide technical expertise in prompt engineering, RAG architectures, model selection, and inference optimization.
- Collaborate with product and engineering teams to communicate customer feedback and influence the platform roadmap.
- Guide customers from proof of concept to production with attention to performance, reliability, and cost efficiency.
Requirements
- At least 5 years of experience in ML/AI systems, including at least 2 years focused on LLMs and generative AI.
- Deep knowledge of the LLM ecosystem, including model architectures and fine-tuning approaches.
- Hands-on experience with prompt engineering, LLM pipeline development and evaluation, agentic frameworks, vector databases, RAG, and LLM-powered applications using OpenAI, Anthropic, or open-source model APIs.
- Strong Python programming skills.
- Excellent communication skills for explaining technical concepts to diverse audiences.
- Experience with vLLM, SGLang, TensorRT-LLM, Transformers, inference optimization, multimodal AI models, Docker, Kubernetes, and open-source ML/AI projects is a bonus.
Benefits
- Remote work from Europe is available.
- Competitive salary and comprehensive benefits package.
- Flexible working arrangements.
- Professional growth opportunities and a dynamic, collaborative work environment.
Tech Stack
Categories
Solutions Engineering
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
