
AI Platform Support Engineer (EMEA)
Lightning AI12 hours ago
London, United KingdomMid Level
Responsibilities
- Partner with customer engineering teams running production training and inference workloads.
- Diagnose distributed systems and ML infrastructure issues and advise customers during incidents and platform degradation events.
- Troubleshoot Kubernetes, GPU, PyTorch, CUDA, NCCL, networking, storage, performance, and scaling issues.
- Analyze logs, metrics, traces, and system behavior to identify root causes.
- Debug containerized workloads across Kubernetes and bare metal GPU environments.
- Drive reliability improvements through incident reviews, tooling, automation, documentation, runbooks, and improved troubleshooting workflows.
- Collaborate with infrastructure, networking, and platform engineering teams.
Requirements
- Strong software engineering and systems troubleshooting background.
- Experience with Kubernetes, containerized environments, cloud infrastructure, and distributed systems.
- Linux knowledge including networking, storage, process management, and performance tuning.
- Experience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry.
- Hands-on experience operating machine learning workloads in production or research environments.
- Experience with distributed ML tooling such as PyTorch, CUDA, or NCCL, plus GPU infrastructure and orchestration.
- Experience troubleshooting performance, reliability, or scaling issues in ML infrastructure.
- Strong communication skills and comfort working directly with technical customers and engineering teams.
- Preferred experience includes large-scale model training or distributed inference, Ray, Kubeflow, Slurm, InfiniBand, RDMA, high-performance networking, bare metal infrastructure, ML storage systems, AI infrastructure or MLOps companies, platform engineering, and Python automation.
Benefits
- Annual base salary range of £75,000—£95,000 GBP, plus eligible bonus and equity.
- Medical, dental, and vision coverage for employees and eligible dependents.
- RSUs, U.K. pension contributions, and location-appropriate retirement benefits.
- Unlimited PTO, company holidays, floating holidays, and a two-week company-wide winter break.
- Paid parental and family leave.
- Annual learning and development allowance.
- Wellness and work-from-home stipends.
- Four weeks of paid sabbatical after four years of service.
- Flexible schedules and a hybrid London work model with at least two in-office days per week.
- Complimentary meals at office hubs.
- Two EMEA shifts are available: Saturday–Tuesday or Thursday–Sunday, 9AM–7PM CET/CEST; occasional team and company offsites may occur.
- Visa sponsorship is not available for this role.
Tech Stack
Categories
Solutions Engineering
About Lightning AI
The AI development platform - From idea to AI, Lightning fast ⚡️. Code together. Prototype. Train on GPUs. Scale. Serve. From your browser - with zero setup. AI Studio is your laptop on the cloud. Zero setup. Always ready. Persistent storage and environments. Code on CPU. Debug on GPU. Scale to multi-node. Run sweeps, jobs and more. Scale models with PyTorch Lightning, Fabric, Lit-GPT, torchmetrics and more.