The Inception Company

Member of Technical Staff, Kernels

The Inception Company
Apply
6 months ago
San Mateo, CA, USASenior

Responsibilities

  • Design and implement custom ML kernels for attention, matrix multiplication, gating, normalization, and other core language-model operations.
  • Develop compute primitives that reduce memory-bandwidth bottlenecks and improve kernel efficiency.
  • Improve infrastructure stability, scalability, reproducibility, precision consistency, and compute utilization.
  • Contribute to distributed compute infrastructure for large-scale language model training and inference.

Requirements

  • BS, MS, or PhD in Computer Science, Engineering, or a related field, or equivalent experience.
  • Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks.
  • Understanding of PyTorch and TensorFlow from a systems perspective.
  • Experience with performance optimization and profiling of ML systems.
  • Experience implementing low-precision formats such as FP8, INT8, or block floating point, or contributing to compiler stacks such as XLA or TVM.
  • Familiarity with data-parallel, model-parallel, and pipeline-parallel distributed training.
  • Proficiency in Python and at least one of C++, Rust, or Go.
  • Experience with Docker, Kubernetes, and CI/CD pipelines.
  • Preferred experience building large-scale language models with tens of billions of parameters or more.
  • Preferred experience with distributed systems and AWS, GCP, or Azure.
  • Preferred familiarity with PyTorch/XLA, DeepSpeed, or Megatron-LM.
  • Preferred open-source contributions to deep learning infrastructure such as PyTorch, DeepSpeed, or XLA.
The Inception Company

About The Inception Company

51-200 employees
Contact me