2 months ago
Remote, United StatesSenior
Base Salary
$170k - $255k/yr
Responsibilities
- Own discovery, technical scoping, infrastructure design, implementation, and production rollout for strategic customer and ISV engagements.
- Build and operate cloud infrastructure and compute orchestration for simulation, training, evaluation, inference, and batch workloads.
- Build platform services for job execution, scheduling, retries, observability, logging, secrets, access control, and cost tracking.
- Create secure, isolated, observable, and reliable onboarding infrastructure including sandbox environments, dataset storage, workflow execution, and deployment.
- Optimize cloud cost, utilization, performance, and reliability and debug failures across application, network, storage, compute, and orchestration layers.
- Partner with Physical AI Systems and Platform & Product FDEs to support GPU-heavy workloads and expose infrastructure through APIs, SDKs, and product workflows.
- Help define infrastructure architecture for multi-tenant SaaS, enterprise deployments, and high-throughput physical AI workloads.
- Convert recurring customer infrastructure problems into reusable platform capabilities and incorporate them into the core platform.
- Use Claude Code, Codex, and Cursor to accelerate production engineering work.
- Co-author reference architectures, solution templates, and technical blogs and maintain customer feedback loops with Field CTO, Product, and Engineering teams.
Requirements
- At least 6 years of hands-on engineering experience in backend, cloud infrastructure, platform engineering, or SRE, including at least 2 years in a customer-facing or deployment-oriented technical role.
- Experience building distributed systems, job orchestration, compute platforms, internal developer platforms, or ML infrastructure.
- Strong Python, Go, or similar systems and backend programming skills.
- Fluency with Claude Code, Codex, and Cursor as AI coding tools used to design, implement, test, debug, and refactor production software.
- Experience with Kubernetes, containers, CI/CD, observability, cloud networking, storage, IAM/RBAC, and infrastructure as code.
- Familiarity with GPU workloads, batch jobs, training pipelines, inference workloads, or HPC-style compute environments.
- Ability to debug infrastructure issues across application, network, storage, compute, and orchestration layers.
- Strong security and reliability instincts involving isolation, RBAC, uptime, and traceability.
- Strong written and verbal communication with customers and internal technical leaders.
- Prior Forward Deployed Engineer or equivalent customer-embedded engineering experience is a bonus.
- Experience with Nebius, AWS, GCP, Azure, Lambda Labs, or other AI cloud infrastructure is a bonus.
- Experience with Slurm, Soperator, Kubernetes GPU scheduling, Ray, Argo, Airflow, Metaflow, or similar orchestration tools is a bonus.
- Experience with ML training infrastructure, model serving, simulation workloads, or large-scale data pipelines is a bonus.
- Experience supporting enterprise customers, design partners, or production pilots is a bonus.
- Familiarity with NVIDIA GPU infrastructure, CUDA workloads, Isaac Sim, Omniverse, or simulation-at-scale is a bonus.
- Applicants must be authorized to work in the country in which they apply and provide proof of employment eligibility.
Benefits
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to 4% company match and immediate vesting.
- 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
- Up to $85/month remote work reimbursement for mobile and internet.
- Company-paid short-term disability, long-term disability, and life insurance.
- Remote work is available from the United States, with the San Francisco Bay Area or Austin, Texas preferred.
- Career growth and learning opportunities, flexibility, ownership, and an international work environment.
Tech Stack
About Nebius
The Nebius AI Cloud brings powerful full-stack infrastructure for AI developers and practitioners across startups, enterprises and science institutes to build and deploy generative AI applications and rapidly deliver scientific breakthroughs by training and running ML models within a secure, high-performance, and cost-optimized cloud environment.