2 months ago
Remote, United StatesSenior
Base Salary
$170k - $255k/yr
Responsibilities
- Own discovery, technical scoping, infrastructure design, implementation, and production rollout for strategic customer and ISV engagements.
- Build and operate cloud infrastructure and compute orchestration for simulation, training, evaluation, inference, and batch workloads.
- Build platform services for job execution, scheduling, retries, observability, logging, secrets, access control, and cost tracking.
- Create secure, isolated, observable, and reliable onboarding infrastructure including sandbox environments, dataset storage, workflow execution, and deployment.
- Optimize cloud cost, utilization, performance, and reliability and debug failures across application, network, storage, compute, and orchestration layers.
- Partner with Physical AI Systems and Platform & Product FDEs to support GPU-heavy workloads and expose infrastructure through APIs, SDKs, and product workflows.
- Help define infrastructure architecture for multi-tenant SaaS, enterprise deployments, and high-throughput physical AI workloads.
- Convert recurring customer infrastructure problems into reusable platform capabilities and incorporate them into the core platform.
- Use Claude Code, Codex, and Cursor to accelerate production engineering work.
- Co-author reference architectures, solution templates, and technical blogs and maintain customer feedback loops with Field CTO, Product, and Engineering teams.
Requirements
- At least 6 years of hands-on engineering experience in backend, cloud infrastructure, platform engineering, or SRE, including at least 2 years in a customer-facing or deployment-oriented technical role.
- Experience building distributed systems, job orchestration, compute platforms, internal developer platforms, or ML infrastructure.
- Strong Python, Go, or similar systems and backend programming skills.
- Fluency with Claude Code, Codex, and Cursor as AI coding tools used to design, implement, test, debug, and refactor production software.
- Experience with Kubernetes, containers, CI/CD, observability, cloud networking, storage, IAM/RBAC, and infrastructure as code.
- Familiarity with GPU workloads, batch jobs, training pipelines, inference workloads, or HPC-style compute environments.
- Ability to debug infrastructure issues across application, network, storage, compute, and orchestration layers.
- Strong security and reliability instincts involving isolation, RBAC, uptime, and traceability.
- Strong written and verbal communication with customers and internal technical leaders.
- Prior Forward Deployed Engineer or equivalent customer-embedded engineering experience is a bonus.
- Experience with Nebius, AWS, GCP, Azure, Lambda Labs, or other AI cloud infrastructure is a bonus.
- Experience with Slurm, Soperator, Kubernetes GPU scheduling, Ray, Argo, Airflow, Metaflow, or similar orchestration tools is a bonus.
- Experience with ML training infrastructure, model serving, simulation workloads, or large-scale data pipelines is a bonus.
- Experience supporting enterprise customers, design partners, or production pilots is a bonus.
- Familiarity with NVIDIA GPU infrastructure, CUDA workloads, Isaac Sim, Omniverse, or simulation-at-scale is a bonus.
- Applicants must be authorized to work in the country in which they apply and provide proof of employment eligibility.
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to 4% company match and immediate vesting.
- 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
- Up to $85/month remote work reimbursement for mobile and internet.
- Company-paid short-term disability, long-term disability, and life insurance.
- Remote work is available from the United States, with the San Francisco Bay Area or Austin, Texas preferred.
- Career growth and learning opportunities, flexibility, ownership, and an international work environment.
Tech Stack
About Nebius
Nebius builds a full-stack AI cloud offering GPU compute, storage, and tools for training and deploying ML models for startups, enterprises, and research labs. It sells consumption-based cloud infrastructure (IaaS/PaaS) and managed services tailored to generative AI workloads, including large-scale model training and inference. The company is headquartered in Amsterdam and operates as an independent provider.
