Lightning AI

Platform Support Engineer

Lightning AI
Apply
3 months ago
Seattle, WA, USA or San Francisco, CA, USASenior

Base Salary

$115k - $140k/yr

Responsibilities

  • Partner directly with customer engineering teams running production training and inference workloads.
  • Diagnose and resolve complex distributed systems and ML infrastructure issues.
  • Advise customers during high-impact incidents and platform degradation events.
  • Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems.
  • Troubleshoot PyTorch, CUDA, NCCL, and inference-serving issues.
  • Analyze logs, metrics, traces, and system behavior to identify root causes.
  • Debug containerized workloads across Kubernetes and bare-metal GPU environments.
  • Support customers scaling workloads across multi-node GPU systems.
  • Diagnose compute, memory, networking, and storage performance bottlenecks.
  • Drive long-term reliability improvements based on recurring customer issues.
  • Contribute to post-incident reviews, automation, internal tooling, documentation, runbooks, observability, and troubleshooting workflows.
  • Collaborate with infrastructure, networking, and platform engineering teams.

Requirements

  • Strong software engineering and systems troubleshooting background.
  • Experience with Kubernetes, containerized environments, cloud infrastructure, and distributed systems.
  • Linux systems knowledge including networking, storage, process management, and performance tuning.
  • Experience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry.
  • Hands-on experience operating machine learning workloads in production or research environments.
  • Experience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL.
  • Familiarity with GPU infrastructure and orchestration.
  • Experience troubleshooting performance, reliability, or scaling issues in ML infrastructure.
  • Understanding of the operational challenges involved in running ML systems at scale.
  • Strong communication skills and the ability to work directly with highly technical customers and engineering teams.
  • Experience with large-scale model training or distributed inference systems is preferred.
  • Familiarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms is preferred.
  • Experience with InfiniBand, RDMA, or high-performance networking is preferred.
  • Experience operating bare-metal infrastructure is preferred.
  • Familiarity with storage systems used in ML environments is preferred.
  • Experience at an AI infrastructure, cloud, MLOps, or developer tooling company is preferred.
  • Contributions to platform engineering, developer infrastructure, or operational tooling projects are preferred.
  • Experience writing automation, tooling, or scripts in Python or similar languages is preferred.

Benefits

  • Hybrid work from the Seattle or San Francisco offices with at least 2 days per week in-office.
  • Monday–Friday schedule with working hours from 8:00 AM to 5:00 PM PST.
  • Occasional team and company offsites.
  • Comprehensive medical, dental, and vision coverage in the U.S.; private medical and dental insurance in the U.K.
  • Retirement and financial wellness support in the U.S.; pension contribution in the U.K.
  • Generous paid time off and holidays.
  • Paid parental leave.
  • Professional development support.
  • Wellness and work-from-home stipends.
  • Flexible work environment.
  • No visa sponsorship is available for this role.

Tech Stack

Categories

DevOpsML EngineeringSolutions Engineering
Lightning AI

About Lightning AI

51-200 employees

The AI development platform - From idea to AI, Lightning fast ⚡️. Code together. Prototype. Train on GPUs. Scale. Serve. From your browser - with zero setup. AI Studio is your laptop on the cloud. Zero setup. Always ready. Persistent storage and environments. Code on CPU. Debug on GPU. Scale to multi-node. Run sweeps, jobs and more. Scale models with PyTorch Lightning, Fabric, Lit-GPT, torchmetrics and more.