2 months ago
Remote, WorldwideMid Level
Responsibilities
- Create technical content demonstrating the use of virtual machines, GPU clusters, Kubernetes, SLURM, Soperator, and other Nebius cloud infrastructure.
- Develop sample code, tutorials, reference architectures, video tutorials, and live coding sessions for cloud computing and ML infrastructure.
- Collaborate with academic partners, including universities, to understand requirements and develop aligned solution architectures.
- Design and document Infrastructure as Code solutions, technical documentation, and how-to guides with the Nebius Solutions Architect Team.
- Advise academic partners on GPU cloud technologies and best practices.
Requirements
- Strong understanding of cloud infrastructure and distributed computing principles.
- Experience with virtual machines, containerization, compute resource management, Infrastructure as Code, and cloud environments.
- Knowledge of GPU clusters, ML workload optimization, networking, storage optimization, resource management, and cloud deployment patterns.
- Working knowledge of Kubernetes and SLURM.
- Strong programming skills in Python and familiarity with the PyTorch ecosystem.
- Excellent written communication skills and ability to clearly explain technical concepts.
- At least 2 years of experience in software development, cloud engineering, DevOps, or a similar technical role.
- Bonus: experience with MLflow, Apache Airflow, Kubeflow, AWS, GCP, Azure ML, NVIDIA NGC, hybrid cloud, on-premises GPU infrastructure, technology partners, or third-party integrations.
- Public presentation skills are a bonus.
Benefits
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
Tech Stack
Categories
Solutions Engineering
About Nebius
The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.