Emerald AI

Member of Technical Staff - AI Cloud Infrastructure

Emerald AI
Apply
2 months ago
Boston, MA, USA +2 moreSenior

Responsibilities

  • Architect managed services from 0 to 1, including GPU capacity productization, isolation boundaries, tenant models, provisioning flows, and scalable service catalogs.
  • Build control-plane services, self-service customer interfaces, automated lifecycle systems, and usage metering integrated with billing infrastructure.
  • Assess and onboard bare-metal GPU infrastructure partners, including evaluating fabric quality, network isolation, economics, and tenant activation workflows.
  • Design and implement secure multi-tenancy across compute, storage, and networking, including InfiniBand and VLAN isolation, quality of service, and encryption despite customer root access.
  • Manage Kubernetes and Slurm environments for large-scale training and inference, including node health, driver fleets, and kernel management across heterogeneous clouds.
  • Deploy and integrate parallel storage systems such as Lustre, VAST, and Weka into the provisioning model.
  • Define service-level objectives, observability standards, and incident-response protocols that align internal standards with provider service-level agreements.

Requirements

  • At least 7 years of infrastructure or platform engineering experience, including architecting and launching a managed cloud or AI platform used by production customers.
  • Strong experience with Kubernetes and Slurm and offering them as managed services.
  • Production experience deploying or operating Lustre or a comparable parallel filesystem such as GPFS, Weka, VAST, or BeeGFS.
  • Strong understanding of cloud service fundamentals, control planes, tenancy and isolation models, APIs, quotas, metering, and operating customer-facing paid services.
  • Deep Linux systems knowledge and mature infrastructure-as-code experience with Terraform and Ansible.
  • Strong programming ability in Python or Go.
  • Familiarity with GPU infrastructure and high-performance networking using InfiniBand, RoCE, and RDMA.
  • Preferred experience at a GPU cloud, hyperscaler AI service, or HPC center delivering compute and storage as a service.
  • Preferred familiarity with NVIDIA SuperPOD, GPUDirect Storage, NCCL debugging, and DCGM.
  • Preferred experience with Lustre multi-tenancy features such as nodemap, fileset mounts, and Kerberos, or with VAST or Weka service-provider deployments.
  • Preferred experience negotiating with and integrating multiple infrastructure vendors and designing for portability.
  • Preferred experience running object storage at scale with S3, Ceph, or MinIO, including data-tiering design.
  • Preferred experience building usage-based billing, metering, or FinOps pipelines.

Benefits

  • Competitive pay and equity, including stock options
  • Medical, dental, vision, and 401(k) matching
  • Flexible location in Washington, D.C., Boston, or the Bay Area
  • Hybrid schedule with 2 work-from-home days per week
  • Opportunity to shape strategy, go-to-market, organizational design, and customer and investor engagement from day one

Categories

Emerald AI

About Emerald AI

11-50 employees

Emerald AI builds the Emerald Conductor software platform for AI data centers, orchestrating training and inference workloads so facilities can dynamically modulate power draw in response to grid conditions while meeting performance targets. It sells enterprise software to data center operators and cloud providers to improve grid reliability and support renewable integration. Founded in 2024 and privately held, it is backed by investors including Radical Ventures and NVIDIA.

Contact me