GrepJob
Liquid AI

Member of Technical Staff - GPU Infrastructure Engineer

Liquid AI
Apply
about 3 hours ago
San Francisco, CA, USAMid Level / Senior
H1B Sponsor

Responsibilities

  • Own the reliability and operation of GPU clusters for training and research.
  • Debug issues across compute, storage, networking, schedulers, and distributed workloads.
  • Improve CPU, GPU, and storage utilization through better tooling and automation.
  • Onboard and migrate workloads across GPU providers and hardware platforms.
  • Build monitoring, validation, and platform abstractions to reduce operational work for researchers.
  • Contribute to the long-term architecture of Liquid AI’s training infrastructure and GPU platform.

Requirements

  • Strong software engineering experience with production-quality infrastructure tooling.
  • Deep knowledge of distributed systems, Linux, networking, and storage.
  • Experience operating a shared compute cluster or distributed training platform.
  • Track record of supporting production users and creating durable solutions.
  • Technical depth to collaborate effectively with senior research and infrastructure engineers.
  • Experience with SLURM, Kubernetes, Ray, Hadoop, or other distributed compute platforms is a plus.
  • Experience supporting GPU, HPC, or large-scale AI training infrastructure is a plus.
  • Familiarity with distributed storage, cluster schedulers, cloud providers, or infrastructure control planes is a plus.

Benefits

  • High-impact ownership of infrastructure affecting AI model training efficiency.
  • Competitive base salary with equity in a unicorn-stage company.
  • 100% coverage of medical, dental, and vision premiums for employees and dependents.
  • 401(k) matching up to 4% of base pay.
  • Unlimited PTO plus company-wide Refill Days throughout the year.

Tech Stack

Apache HadoopKubernetesLinux

Categories

AI & MLData EngineeringDevOps