
Lightning AI
Lightning AI builds an end-to-end platform for prototyping, training, and deploying AI systems, alongside open-source tools such as PyTorch Lightning and torchmetrics. It offers a browser-based studio and managed GPU infrastructure for teams ranging from solo researchers to large enterprises, with built-in observability and controls. Founded in 2019 and headquartered in New York, the company merged with Voltage Park to pair developer-first software with large-scale, cost-efficient compute.
Open Positions at Lightning AI
15 open positions
Senior Software Engineer building backend and distributed systems for Lightning AI’s core AI platform. The role combines cloud integrations, infrastructure automation, reliability, monitoring, and technical mentorship.
Senior Software Engineer building backend and distributed systems for Lightning AI’s cloud platform, including APIs, infrastructure automation, and resource management services. The role offers hybrid work near Redmond, Washington, with a focus on scalable AI infrastructure and production reliability.
Build and scale the user interface and frontend infrastructure for Lightning AI’s platform, delivering high-impact features used by thousands of organizations. You’ll work cross-functionally, influence technical direction, and help improve frontend quality and performance.
AI Platform Support Engineer partnering directly with ML engineering teams across APAC to diagnose complex production issues in distributed training, inference, Kubernetes, GPU, and cloud infrastructure. The remote role combines deep systems troubleshooting with reliability improvements, automation, and customer-facing technical guidance.
Build production software, APIs, and automation that operate Lightning AI’s large-scale GPU, bare-metal, storage, and networking infrastructure. You’ll improve provisioning, lifecycle management, observability, and reliability across thousands of servers supporting AI/ML and HPC workloads.
Build and optimize large language model training and post-training systems at Lightning AI, improving model quality, efficiency, and production readiness. The role combines frontier model research, PyTorch engineering, distributed systems, and customer-informed platform development.
Build the backend infrastructure, APIs, and orchestration systems that power large-scale AI training, experimentation, and agent workflows. You’ll help shape the developer experience used by researchers, startups, and enterprise AI teams worldwide.
Build the backend infrastructure that makes production-ready AI agent applications reliable, scalable, and easy for developers to create. You’ll work on orchestration, tool execution, workflows, memory, state, and developer-facing APIs while helping shape a growing platform.
Build the backend control planes and automation that turn Lightning AI’s large-scale GPU capacity into reliable, production-ready Kubernetes and Slurm environments. This senior role combines distributed systems, cloud infrastructure, and hands-on platform operations.
Help ML engineering teams diagnose and scale complex training and inference workloads across Kubernetes, GPUs, and distributed systems. This customer-facing engineering role combines deep infrastructure troubleshooting with platform reliability improvements.
Unlock 5 More Matching Jobs
Create an account to view all jobs matching your criteria.
First to know
Discover the latest jobs before everyone else
Zero Spam
No ghost jobs, reposts, or sponsored listings
AI-powered filters
Find the most relevant jobs for you