Liquid AI

Member of Technical Staff - GPU Infrastructure Engineer

Liquid AI
Apply
2 months ago

Responsibilities

  • Own the reliability and operation of GPU clusters used for training and research
  • Debug issues across compute, storage, networking, schedulers, and distributed workloads
  • Improve CPU, GPU, and storage utilization through tooling and automation
  • Onboard and migrate workloads across GPU providers and hardware platforms
  • Build monitoring, validation, and platform abstractions that reduce operational work for researchers
  • Contribute to the architecture of the training infrastructure and GPU platform
  • Respond to immediate operational issues while replacing manual work with durable automation

Requirements

  • Strong software engineering experience building production-quality infrastructure tooling and automation
  • Deep knowledge of distributed systems, Linux, networking, and storage
  • Experience operating a shared compute cluster or distributed training platform
  • Experience supporting production users and converting recurring failures into durable solutions
  • Technical depth to partner effectively with senior research and infrastructure engineers
  • Nice to have experience with SLURM, Kubernetes, Ray, Hadoop, or another distributed compute platform
  • Nice to have experience supporting GPU, HPC, or large-scale AI training infrastructure
  • Nice to have experience with distributed storage, cluster schedulers, cloud providers, or infrastructure control planes

Benefits

  • 100% of medical, dental, and vision premiums paid for employees and dependents
  • 401(k) matching up to 4% of base pay
  • Unlimited PTO and company-wide Refill Days
  • Competitive base salary with equity
  • Own high-impact infrastructure supporting foundation model training

Tech Stack

Apache HadoopKubernetesLinux

Categories

DevOpsSite Reliability
Liquid AI

About Liquid AI

51-200 employees

Liquid AI builds general-purpose AI systems that run efficiently from data center accelerators to on-device hardware, emphasizing low latency, memory efficiency, privacy, and reliability. The company partners with enterprises in consumer electronics, automotive, life sciences, and financial services to deploy and benchmark models for real-world workloads. Founded in 2023 out of MIT CSAIL and headquartered in Cambridge, Massachusetts, it is privately held.

Contact me