
Platform Support Engineer
Lightning AI3 months ago
Seattle, WA, USA or San Francisco, CA, USASenior
Base Salary
$115k - $140k/yr
Responsibilities
- Partner directly with customer engineering teams running production training and inference workloads.
- Diagnose and resolve complex distributed systems and ML infrastructure issues.
- Advise customers during high-impact incidents and platform degradation events.
- Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems.
- Troubleshoot PyTorch, CUDA, NCCL, and inference-serving issues.
- Analyze logs, metrics, traces, and system behavior to identify root causes.
- Debug containerized workloads across Kubernetes and bare-metal GPU environments.
- Support customers scaling workloads across multi-node GPU systems.
- Diagnose compute, memory, networking, and storage performance bottlenecks.
- Drive long-term reliability improvements based on recurring customer issues.
- Contribute to post-incident reviews, automation, internal tooling, documentation, runbooks, observability, and troubleshooting workflows.
- Collaborate with infrastructure, networking, and platform engineering teams.
Requirements
- Strong software engineering and systems troubleshooting background.
- Experience with Kubernetes, containerized environments, cloud infrastructure, and distributed systems.
- Linux systems knowledge including networking, storage, process management, and performance tuning.
- Experience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry.
- Hands-on experience operating machine learning workloads in production or research environments.
- Experience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL.
- Familiarity with GPU infrastructure and orchestration.
- Experience troubleshooting performance, reliability, or scaling issues in ML infrastructure.
- Understanding of the operational challenges involved in running ML systems at scale.
- Strong communication skills and the ability to work directly with highly technical customers and engineering teams.
- Experience with large-scale model training or distributed inference systems is preferred.
- Familiarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms is preferred.
- Experience with InfiniBand, RDMA, or high-performance networking is preferred.
- Experience operating bare-metal infrastructure is preferred.
- Familiarity with storage systems used in ML environments is preferred.
- Experience at an AI infrastructure, cloud, MLOps, or developer tooling company is preferred.
- Contributions to platform engineering, developer infrastructure, or operational tooling projects are preferred.
- Experience writing automation, tooling, or scripts in Python or similar languages is preferred.
Benefits
- Hybrid work from the Seattle or San Francisco offices with at least 2 days per week in-office.
- Monday–Friday schedule with working hours from 8:00 AM to 5:00 PM PST.
- Occasional team and company offsites.
- Comprehensive medical, dental, and vision coverage in the U.S.; private medical and dental insurance in the U.K.
- Retirement and financial wellness support in the U.S.; pension contribution in the U.K.
- Generous paid time off and holidays.
- Paid parental leave.
- Professional development support.
- Wellness and work-from-home stipends.
- Flexible work environment.
- No visa sponsorship is available for this role.
Tech Stack
Categories
About Lightning AI
The AI development platform - From idea to AI, Lightning fast ⚡️. Code together. Prototype. Train on GPUs. Scale. Serve. From your browser - with zero setup. AI Studio is your laptop on the cloud. Zero setup. Always ready. Persistent storage and environments. Code on CPU. Debug on GPU. Scale to multi-node. Run sweeps, jobs and more. Scale models with PyTorch Lightning, Fabric, Lit-GPT, torchmetrics and more.