Lightning AI

AI Platform Support Engineer (EMEA)

Lightning AI
Apply
12 hours ago
London, United KingdomMid Level

Responsibilities

  • Partner with customer engineering teams running production training and inference workloads.
  • Diagnose distributed systems and ML infrastructure issues and advise customers during incidents and platform degradation events.
  • Troubleshoot Kubernetes, GPU, PyTorch, CUDA, NCCL, networking, storage, performance, and scaling issues.
  • Analyze logs, metrics, traces, and system behavior to identify root causes.
  • Debug containerized workloads across Kubernetes and bare metal GPU environments.
  • Drive reliability improvements through incident reviews, tooling, automation, documentation, runbooks, and improved troubleshooting workflows.
  • Collaborate with infrastructure, networking, and platform engineering teams.

Requirements

  • Strong software engineering and systems troubleshooting background.
  • Experience with Kubernetes, containerized environments, cloud infrastructure, and distributed systems.
  • Linux knowledge including networking, storage, process management, and performance tuning.
  • Experience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry.
  • Hands-on experience operating machine learning workloads in production or research environments.
  • Experience with distributed ML tooling such as PyTorch, CUDA, or NCCL, plus GPU infrastructure and orchestration.
  • Experience troubleshooting performance, reliability, or scaling issues in ML infrastructure.
  • Strong communication skills and comfort working directly with technical customers and engineering teams.
  • Preferred experience includes large-scale model training or distributed inference, Ray, Kubeflow, Slurm, InfiniBand, RDMA, high-performance networking, bare metal infrastructure, ML storage systems, AI infrastructure or MLOps companies, platform engineering, and Python automation.

Benefits

  • Annual base salary range of £75,000—£95,000 GBP, plus eligible bonus and equity.
  • Medical, dental, and vision coverage for employees and eligible dependents.
  • RSUs, U.K. pension contributions, and location-appropriate retirement benefits.
  • Unlimited PTO, company holidays, floating holidays, and a two-week company-wide winter break.
  • Paid parental and family leave.
  • Annual learning and development allowance.
  • Wellness and work-from-home stipends.
  • Four weeks of paid sabbatical after four years of service.
  • Flexible schedules and a hybrid London work model with at least two in-office days per week.
  • Complimentary meals at office hubs.
  • Two EMEA shifts are available: Saturday–Tuesday or Thursday–Sunday, 9AM–7PM CET/CEST; occasional team and company offsites may occur.
  • Visa sponsorship is not available for this role.

Tech Stack

Categories

Solutions Engineering
Lightning AI

About Lightning AI

51-200 employees

The AI development platform - From idea to AI, Lightning fast ⚡️. Code together. Prototype. Train on GPUs. Scale. Serve. From your browser - with zero setup. AI Studio is your laptop on the cloud. Zero setup. Always ready. Persistent storage and environments. Code on CPU. Debug on GPU. Scale to multi-node. Run sweeps, jobs and more. Scale models with PyTorch Lightning, Fabric, Lit-GPT, torchmetrics and more.

Contact me