
Staff Software Engineer - GenAI inference
Databricks11 months ago
Base Salary
$191k - $233k/yr
Responsibilities
- Own the architecture, design, and implementation of the GenAI inference engine and collaborate on an LLM-optimized model-serving stack.
- Partner with researchers to integrate new model architectures and features such as sparsity, activation compression, and mixture-of-experts.
- Lead optimization of latency, throughput, memory efficiency, and hardware utilization across GPUs and accelerators.
- Define standards and guide instrumentation, profiling, and tracing tooling to identify bottlenecks.
- Architect routing, batching, scheduling, memory management, and dynamic-loading mechanisms for inference workloads.
- Ensure reliability, reproducibility, and fault tolerance through A/B launches, rollback, and model versioning.
- Integrate with federated and distributed inference infrastructure across nodes, including load balancing and communication management.
- Collaborate with platform engineering, cloud infrastructure, and security/compliance teams.
- Represent the team through benchmarks, whitepapers, and open-source contributions.
Requirements
- BS, MS, or PhD in Computer Science or a related field, or equivalent experience.
- 6+ years or equivalent experience in performance-critical software systems.
- Strong ownership of complex system components and end-to-end architectural decision-making.
- Deep understanding of ML inference internals, including attention, MLPs, recurrent modules, quantization, and sparse operations.
- Hands-on CUDA, GPU programming, and experience with cuBLAS, cuDNN, and NCCL.
- Strong distributed-systems design background, including RPC frameworks, queuing, RPC batching, sharding, and memory partitioning.
- Experience identifying and resolving performance bottlenecks across kernels, memory, networking, and scheduling layers.
- Experience building instrumentation, tracing, and profiling tools for ML models.
- Ability to lead through influence and collaborate closely with ML researchers to productionize new model ideas.
- Strong communication, leadership, proactive ownership, and collaboration skills.
- Published research or open-source contributions in ML systems, inference optimization, or model serving are a bonus.
Tech Stack
Apache SparkMLflow
Categories
About Databricks
Databricks builds a cloud-based data and AI platform centered on the lakehouse architecture, combining data engineering, analytics, and machine learning with Apache Spark, Delta Lake, and MLflow. It sells subscriptions and cloud services to enterprises that need to unify data pipelines and develop large-scale AI and analytics. Founded in 2013 by the creators of Apache Spark and headquartered in San Francisco, the company is privately held and serves organizations across many industries.