8 hours ago
Remote, United States or Remote, EMEASenior
Responsibilities
- Own discovery, technical scoping, infrastructure design, implementation, and production rollout for strategic customer and ISV engagements.
- Build and operate cloud infrastructure and compute orchestration for simulation, training, evaluation, inference, and batch workloads.
- Develop platform services for job execution, scheduling, retries, observability, logging, secrets, access control, and cost tracking.
- Build secure, isolated, observable onboarding infrastructure for pilots, including sandbox environments, dataset storage, workflow execution, and deployment.
- Optimize reliability, performance, utilization, and cloud cost while debugging issues across application, network, storage, compute, and orchestration layers.
- Partner with Physical AI Systems and Platform & Product FDEs to support GPU workloads and expose infrastructure capabilities through APIs, SDKs, and product workflows.
- Define infrastructure architecture for multi-tenant SaaS, enterprise deployments, and high-throughput physical AI workloads.
- Convert recurring customer infrastructure problems into reusable platform capabilities and contribute feedback to Product, Engineering, and the Physical AI roadmap.
- Use Claude Code, Codex, and Cursor to accelerate the design, implementation, testing, debugging, and refactoring of production software.
- Create reference architectures, solution templates, technical blogs, and structured customer feedback loops for the broader field organization.
Requirements
- At least 6 years of hands-on backend, cloud infrastructure, platform engineering, or SRE experience, including at least 2 years in a customer-facing or deployment-oriented technical role.
- Experience building distributed systems, job orchestration, compute platforms, internal developer platforms, or ML infrastructure.
- Strong Python, Go, or comparable systems and backend programming skills.
- Fluency with Claude Code, Codex, and Cursor as part of an AI-native development workflow.
- Experience with Kubernetes, cloud-native tooling, observability, cloud networking, storage, IAM/RBAC, and infrastructure as code.
- Familiarity with GPU workloads, batch jobs, training pipelines, inference workloads, or HPC-style computing environments.
- Ability to debug infrastructure issues across application, network, storage, compute, and orchestration layers.
- Strong security and reliability instincts involving isolation, RBAC, uptime, and traceability for customer workloads.
- Ability to operate with high autonomy and ambiguity while designing simple, composable infrastructure.
- Strong written and verbal communication skills for working with customer CTOs, design partners, and internal leadership.
- Prior Forward Deployed Engineer or equivalent customer-embedded engineering experience is a plus.
- Experience with Nebius, AWS, GCP, Azure, Lambda Labs, Slurm, Soperator, Kubernetes GPU scheduling, Ray, Argo, Airflow, Metaflow, ML training infrastructure, model serving, simulation workloads, large-scale data pipelines, enterprise customers, design partners, or production pilots is a plus.
- Familiarity with NVIDIA GPU infrastructure, CUDA workloads, Isaac Sim, Omniverse, or simulation at scale is a plus.
Benefits
- Competitive compensation, career growth, and learning opportunities.
- Flexibility and ownership in a collaborative, innovative, and international environment.
- Opportunity to work on impactful AI projects with talented teams.
- Remote work is available from the United States, with the SF Bay Area or Austin, Texas preferred.
Tech Stack
Categories
Forward Deployed
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
