
Staff Software Engineer - AI Research Infrastructure
Databricks5 months ago
Base Salary
$190k - $270k/yr
Responsibilities
- Design and implement infrastructure for large-scale experiments, data processing, and model training across HPC clusters, GPU fleets, and cloud-based systems.
- Build abstractions for job submission, scheduling, and monitoring that enable researchers to launch large-scale experiments quickly.
- Create experiment management systems, CI/testing infrastructure for research code, and other developer productivity tooling.
- Influence the roadmap for how AI Research trains, evaluates, and ships models.
- Mentor and serve as a technical force multiplier for engineers working on compute, infrastructure, and AI systems.
- Partner with research scientists, ML engineers, and platform teams to turn experimental workloads into robust, repeatable pipelines.
Requirements
- BS, MS, or PhD in Computer Science or a related field.
- At least 5 years of software engineering experience, including substantial experience with large-scale distributed systems or infrastructure.
- Deep experience building and operating distributed systems, data pipelines, or large-scale backend services, ideally involving GPUs, clusters, or major cloud providers.
- Proficiency in one or more systems programming languages such as C++, Rust, Go, Java, or Scala.
- Experience building or significantly contributing to cluster schedulers, resource managers, or large-scale job orchestration systems such as Kubernetes, Slurm, Ray, or custom internal systems.
- Understanding of modern ML training and inference workflows, including distributed training, model parallelism, fine-tuning, and evaluation.
- Experience driving complex systems from prototype to stable, well-owned services with strong operational practices.
- Clear communication skills and the ability to translate between research needs and infrastructure realities.
Benefits
- Comprehensive benefits and perks are offered, with details varying by region.
Tech Stack
About Databricks
Databricks builds a cloud-based data and AI platform centered on the lakehouse architecture, combining data engineering, analytics, and machine learning with Apache Spark, Delta Lake, and MLflow. It sells subscriptions and cloud services to enterprises that need to unify data pipelines and develop large-scale AI and analytics. Founded in 2013 by the creators of Apache Spark and headquartered in San Francisco, the company is privately held and serves organizations across many industries.