Databricks

Staff Software Engineer - AI Research Infrastructure

Databricks
Apply
5 months ago
San Francisco, CA, USA or New York, NY, USAStaff+
H1B sponsor

Base Salary

$190k - $270k/yr

Responsibilities

  • Design and implement infrastructure for large-scale experiments, data processing, and model training across HPC clusters, GPU fleets, or cloud-based systems.
  • Build abstractions for job submission, scheduling, and monitoring that enable rapid research experimentation.
  • Create experiment management systems, CI and testing infrastructure for research code, and workflows that improve developer productivity.
  • Influence the roadmap for how AI Research trains, evaluates, and ships models.
  • Mentor and serve as a technical force multiplier for engineers working on compute, infrastructure, and AI systems.
  • Partner with research scientists, ML engineers, and platform teams to develop robust, repeatable pipelines.

Requirements

  • BS, MS, or PhD in Computer Science or a related field.
  • 5+ years of software engineering experience, including substantial experience with large-scale distributed systems or infrastructure.
  • Deep experience building and operating distributed systems, data pipelines, or large-scale backend services, ideally involving GPUs, clusters, or major cloud providers.
  • Proficiency in one or more systems programming languages such as C++, Rust, Go, Java, or Scala.
  • Experience building or significantly contributing to cluster schedulers, resource managers, or large-scale job orchestration systems such as Kubernetes, Slurm, or Ray.
  • Understanding of modern ML training and inference workflows, including distributed training, model parallelism, fine-tuning, and evaluation.
  • Experience driving complex systems from prototype to stable, well-owned services with strong operational practices.
  • Clear communication with both researchers and engineers.
Databricks

About Databricks

10,000+ employees

Databricks builds a cloud-based data and AI platform centered on the lakehouse architecture, combining data engineering, analytics, and machine learning with Apache Spark, Delta Lake, and MLflow. It sells subscriptions and cloud services to enterprises that need to unify data pipelines and develop large-scale AI and analytics. Founded in 2013 by the creators of Apache Spark and headquartered in San Francisco, the company is privately held and serves organizations across many industries.

Contact me