3 months ago
Remote, EMEAMid Level
Responsibilities
- Create technical content demonstrating the use of virtual machines, GPU clusters, Kubernetes, SLURM, Soperator, and other Nebius cloud infrastructure.
- Develop sample code, tutorials, reference architectures, video tutorials, and live coding sessions for cloud computing and ML infrastructure.
- Collaborate with academic partners, including universities, to understand requirements and develop aligned solution architectures.
- Design and document Infrastructure as Code solutions, technical documentation, and how-to guides with the Nebius Solutions Architect Team.
- Advise academic partners on GPU cloud technologies and best practices.
Requirements
- Strong understanding of cloud infrastructure and distributed computing principles.
- Experience with virtual machines, containerization, compute resource management, Infrastructure as Code, and cloud environments.
- Knowledge of GPU clusters, ML workload optimization, networking, storage optimization, resource management, and cloud deployment patterns.
- Working knowledge of Kubernetes and SLURM.
- Strong programming skills in Python and familiarity with the PyTorch ecosystem.
- Excellent written communication skills and ability to clearly explain technical concepts.
- At least 2 years of experience in software development, cloud engineering, DevOps, or a similar technical role.
- Bonus: experience with MLflow, Apache Airflow, Kubeflow, AWS, GCP, Azure ML, NVIDIA NGC, hybrid cloud, on-premises GPU infrastructure, technology partners, or third-party integrations.
- Public presentation skills are a bonus.
Benefits
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
Tech Stack
Categories
Solutions Engineering
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
