
Lightning AI
The AI development platform - From idea to AI, Lightning fast ⚡️. Code together. Prototype. Train on GPUs. Scale. Serve. From your browser - with zero setup. AI Studio is your laptop on the cloud. Zero setup. Always ready. Persistent storage and environments. Code on CPU. Debug on GPU. Scale to multi-node. Run sweeps, jobs and more. Scale models with PyTorch Lightning, Fabric, Lit-GPT, torchmetrics and more.
Open Positions at Lightning AI
11 open positions
Build and optimize large language model training and post-training systems at Lightning AI, improving model quality, efficiency, and production readiness. The role combines frontier model research, PyTorch engineering, distributed systems, and customer-informed platform development.
Build the backend infrastructure, APIs, and orchestration systems that power large-scale AI training, experimentation, and agent workflows. You’ll help shape the developer experience used by researchers, startups, and enterprise AI teams worldwide.
Build and scale the foundational backend platform powering Lightning AI’s authentication, APIs, multi-tenancy, resource management, and developer workflows. This senior role combines distributed-systems engineering, platform reliability, technical leadership, and mentorship.
Build the backend infrastructure that makes production-ready AI agent applications reliable, scalable, and easy for developers to create. You’ll work on orchestration, tool execution, workflows, memory, state, and developer-facing APIs while helping shape a growing platform.
Build the backend control planes and automation that turn Lightning AI’s large-scale GPU capacity into reliable, production-ready Kubernetes and Slurm environments. This senior role combines distributed systems, cloud infrastructure, and hands-on platform operations.
Help ML engineering teams diagnose and scale complex training and inference workloads across Kubernetes, GPUs, and distributed systems. This customer-facing engineering role combines deep infrastructure troubleshooting with platform reliability improvements.
Build and deploy production AI systems directly with customers, taking engagements from ambiguous business goals through scalable, observable production delivery. This hands-on role combines software engineering, AI infrastructure, technical customer partnership, and product thinking.
Build and scale Lightning AI’s platform across React-based frontend systems, APIs, CLI tools, and backend services. Own features end-to-end while improving architecture, reliability, delivery speed, and engineering practices for machine learning workloads.
Build and scale Lightning AI’s backend platform in Go, spanning APIs, infrastructure, billing, security, and integrations. You’ll own features end-to-end, shape technical direction, and help deliver reliable systems for machine learning workloads.
Optimize deep learning training and inference across compilers, kernels, and distributed systems as a Research Engineer at Lightning AI. You’ll advance the Thunder compiler and PyTorch Lightning ecosystem to improve model performance and efficiency at scale.
Unlock 1 More Matching Jobs
Choose a plan to access all jobs matching your criteria.
First to know
Discover the latest jobs before everyone else
Zero Spam
No ghost jobs, reposts, or sponsored listings
AI-powered filters
Find the most relevant jobs for you