6 months ago
Remote, United States +2 moreSenior
Responsibilities
- Build prototypes and demos across serverless inference, databases, MLflow, MLOps, Physical AI, HCLS, and other applied AI use cases.
- Support customers through proof-of-concept design, technical onboarding, validation, and hands-on ML stack integration.
- Research emerging training techniques, inference optimizations, agentic architectures, and frameworks, then turn findings into prototypes, writeups, and product recommendations.
- Provide specific customer-informed feedback to the product roadmap and identify changes needed based on POC experience.
- Create reusable notebooks, reference architectures, benchmark results, and other technical assets that reduce onboarding friction.
- Develop a library of polished demos for sales, product, and engineering teams and help customers achieve faster time-to-value.
Requirements
- Hands-on experience fine-tuning large models, debugging distributed training jobs, building production RAG or agentic pipelines, and optimizing inference on GPU infrastructure.
- Fluency with PyTorch, HuggingFace, CUDA fundamentals, Kubernetes for ML, MLflow or an equivalent tool, and vector databases.
- Experience working with enterprise ML teams as a solutions engineer, customer engineer, or closely customer-facing ML engineer.
- Ability to read research papers and implement the techniques in working systems.
- Ability to explain technical ML tradeoffs to engineers and cost implications to CTO-level stakeholders.
Benefits
- Competitive salary and comprehensive benefits package.
- Opportunities for professional growth within Nebius.
- Flexible working arrangements.
- Dynamic and collaborative work environment.
Tech Stack
Categories
Solutions Engineering
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
