Staff Software Engineer - GenAI Performance and Kernel
Databricks11 months ago
Base Salary
$191k - $233k/yr
Responsibilities
- Lead the design, implementation, benchmarking, and maintenance of compute kernels such as attention, MLP, softmax, and layer normalization across GPU and accelerator backends.
- Drive the kernel performance roadmap through vectorization, tensorization, tiling, fusion, mixed precision, sparsity, quantization, memory reuse, scheduling, and auto-tuning.
- Integrate kernel optimizations with higher-level ML systems and production inference pipelines.
- Build profiling, instrumentation, and verification tooling to identify correctness issues, performance regressions, numerical problems, and hardware utilization gaps.
- Lead root-cause investigations into inference bottlenecks such as memory bandwidth, cache contention, kernel launch overhead, and tensor fragmentation.
- Create reusable abstractions and frameworks that support modularity, cross-backend portability, and maintainability.
- Influence system architecture decisions involving memory layout, dataflow scheduling, and kernel fusion boundaries.
- Mentor engineers, conduct code reviews, establish best practices, and collaborate with ML, infrastructure, tooling, and product teams.
- Roll out kernel-level optimizations into production and monitor their impact.
Requirements
- Bachelor’s, master’s, or PhD in Computer Science or a related field.
- Deep hands-on experience writing and tuning compute kernels for ML workloads using technologies such as CUDA, Triton, OpenCL, LLVM IR, or assembly.
- Strong knowledge of GPU and accelerator architecture, including memory hierarchy, tensor cores, scheduling, and SM occupancy.
- Experience with optimization techniques including tiling, blocking, software pipelining, vectorization, fusion, loop transformations, and auto-tuning.
- Familiarity with ML kernel libraries such as cuBLAS, cuDNN, CUTLASS, oneDNN, or open kernels.
- Strong debugging and profiling skills using tools such as Nsight, NVProf, perf, vtune, or custom instrumentation.
- Experience with numerical stability, mixed precision, quantization, error propagation, and integration of optimized kernels into real-world ML inference systems.
- Experience with distributed inference pipelines, memory management, runtime systems, and GPU-accelerated products.
- Track record of shipping performance-critical, high-quality production software.
- Published work in systems or ML performance venues, experience with custom accelerators or FPGA, and experience with sparsity or model compression are bonuses.
Benefits
- Comprehensive employee benefits and perks, with region-specific details available through the Databricks benefits site.
- Annual performance bonus and equity may be included in the total compensation package.
Tech Stack
Categories
About Databricks
Databricks is the Data and AI company. More than 20,000 organizations worldwide — including adidas, AT&T, Bayer, Block, Mastercard, Rivian, Unilever, and over 60% of the Fortune 500 — rely on Databricks to build and scale data and AI apps, analytics and agents. Headquartered in San Francisco with 30+ offices around the globe, Databricks offers a unified Data Intelligence Platform that includes Agent Bricks, Lakeflow, Lakehouse, Lakebase and Unity Catalog. --- Databricks applicants Please apply through our official Careers page at databricks.com/company/careers. All official communication from Databricks will come from email addresses ending with @databricks.com or @goodtime.io (our meeting tool).