2 months ago
Remote, EMEAMid Level
Responsibilities
- Create technical educational content demonstrating workloads using virtual machines, GPU clusters, Kubernetes, SLURM, and Soperator.
- Develop sample code, tutorials, reference architectures, video tutorials, and live coding sessions for cloud computing and ML infrastructure.
- Collaborate with academic partners to understand requirements and develop aligned solution architectures.
- Design and document Infrastructure as Code solutions, technical documentation, and how-to guides with the Solutions Architect team.
- Advise academic partners on GPU cloud technologies, infrastructure, and best practices.
Requirements
- Strong understanding of cloud infrastructure and distributed computing principles, including virtual machines, containerization, compute resources, networking, storage optimization, and resource management.
- Experience with Infrastructure as Code, preferably Terraform, and prior work with containerization and cloud environments.
- Knowledge of GPU clusters and techniques for optimizing diverse ML and cloud workloads.
- Working knowledge of Kubernetes and SLURM.
- Strong programming skills in Python and familiarity with the PyTorch ecosystem.
- Excellent written communication skills and the ability to express technical ideas clearly in text.
- At least 2 years of experience in software development, cloud engineering, DevOps, or a similar technical role.
- Experience with cloud technologies and infrastructure is required.
- Bonus qualifications include MLflow, Apache Airflow, Kubeflow, AWS, GCP, Azure ML, NVIDIA NGC, hybrid or on-premises GPU infrastructure, technology partnerships, third-party integrations, and public presentations.
Benefits
- Competitive compensation.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment with talented teams.
Tech Stack
Categories
Solutions Engineering
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
