Member of Technical Staff, Kernels
The Inception Company6 months ago
San Mateo, CA, USASenior
Responsibilities
- Design and implement custom ML kernels for attention, matrix multiplication, gating, normalization, and other core language-model operations.
- Develop compute primitives that reduce memory-bandwidth bottlenecks and improve kernel efficiency.
- Improve infrastructure stability, scalability, reproducibility, precision consistency, and compute utilization.
- Contribute to distributed compute infrastructure for large-scale language model training and inference.
Requirements
- BS, MS, or PhD in Computer Science, Engineering, or a related field, or equivalent experience.
- Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks.
- Understanding of PyTorch and TensorFlow from a systems perspective.
- Experience with performance optimization and profiling of ML systems.
- Experience implementing low-precision formats such as FP8, INT8, or block floating point, or contributing to compiler stacks such as XLA or TVM.
- Familiarity with data-parallel, model-parallel, and pipeline-parallel distributed training.
- Proficiency in Python and at least one of C++, Rust, or Go.
- Experience with Docker, Kubernetes, and CI/CD pipelines.
- Preferred experience building large-scale language models with tens of billions of parameters or more.
- Preferred experience with distributed systems and AWS, GCP, or Azure.
- Preferred familiarity with PyTorch/XLA, DeepSpeed, or Megatron-LM.
- Preferred open-source contributions to deep learning infrastructure such as PyTorch, DeepSpeed, or XLA.