
Member of Technical Staff - AI Cloud Infrastructure
Emerald AIResponsibilities
- Architect and productize managed GPU services, including isolation boundaries, tenant models, provisioning flows, and service catalogs across diverse providers.
- Build control-plane services, self-service customer interfaces, automated lifecycle systems, and usage metering integrated with billing infrastructure.
- Evaluate and onboard bare-metal GPU infrastructure partners, including their fabric quality, network isolation, economics, and tenant handoff processes.
- Design and implement secure multi-tenancy across compute, storage, and networking, including InfiniBand and VLANs, QoS, encryption, and root-access scenarios.
- Manage Kubernetes and Slurm environments for large-scale training and inference, including node health, driver fleets, and kernel management across heterogeneous clouds.
- Deploy and integrate parallel storage systems such as Lustre, VAST, and Weka into the provisioning model.
- Define SLOs, observability standards, and incident-response protocols aligned with provider SLAs.
Requirements
- At least 7 years of infrastructure or platform engineering experience, including architecting and launching a managed cloud or AI platform used by production customers.
- Strong production experience with Kubernetes and Slurm as managed services.
- Production experience deploying or operating Lustre or comparable parallel filesystems such as GPFS, Weka, VAST, or BeeGFS, including knowledge of architecture, tuning, and failure modes.
- Strong understanding of cloud service fundamentals, including control planes, tenancy and isolation models, APIs, quotas, metering, and operating paid services.
- Deep Linux systems knowledge, mature infrastructure-as-code experience with Terraform and Ansible, and programming ability in Python or Go.
- Familiarity with GPU infrastructure and high-performance networking using InfiniBand, RoCE, and RDMA, as well as the GPU software stack.
- Preferred experience at a GPU cloud, hyperscaler AI service, or HPC center providing compute and storage as a service.
- Preferred familiarity with NVIDIA SuperPOD reference architectures, GPUDirect Storage, NCCL debugging, and DCGM.
- Preferred experience with Lustre multi-tenancy features such as nodemap, fileset mounts, and Kerberos, or with VAST or Weka service-provider deployments.
- Preferred experience integrating multiple infrastructure vendors, designing for portability, operating object storage at scale, and building billing, metering, or FinOps pipelines.
Benefits
- Competitive pay and equity, including stock options.
- Medical, dental, vision, and 401(k) matching.
- Flexible location from Washington, D.C., Boston, or the Bay Area, with two work-from-home days per week.
- Opportunity to shape strategy, go-to-market, organizational design, and customer and investor engagement from day one.
About Emerald AI
Emerald AI transforms energy-intensive data centers into AI-powered grid allies. The Emerald AI Conductor platform enables AI data centers to flexibly adjust their power consumption from the electricity grid on demand, orchestrating the computing loads of inference, training, and fine-tuning AI models across a network of data centers to bolster the power grid’s reliability while meeting strict compute performance standards. Flexibility could unlock up to 100 GW, or more than a decade of U.S. AI data center growth and trillions in AI investment, without enormous investments in electric grids or power plants—advancing innovation and competitiveness, ensuring power affordability for communities across the country, and bolstering the reliability of the electric power system.