GrepJob
Fireworks AI

Member of Technical Staff, Performance Optimization

Fireworks AI
Apply
over 1 year ago
San Mateo, CA, USAStaff+
H1B Sponsor

Base Salary

$175k - $220k/yr

Responsibilities

  • Optimize system and GPU performance for high-throughput AI workloads across training and inference.
  • Analyze and improve latency, throughput, memory usage, and compute efficiency.
  • Profile systems to identify and resolve GPU- and kernel-level bottlenecks.
  • Implement low-level optimizations using CUDA, Triton, and other performance tooling.
  • Improve execution speed and resource utilization for LLM, VLM, and video-model workloads.
  • Collaborate with ML researchers to co-design and tune model architectures for hardware efficiency.
  • Improve support for mixed precision, quantization, and model graph optimization.
  • Build and maintain performance benchmarking and monitoring infrastructure.
  • Scale inference and training systems across multi-GPU and multi-node environments.
  • Evaluate and integrate optimizations for emerging hardware accelerators and specialized runtimes.
  • Develop asynchronous low-latency sampling, GPU kernels, distributed routing, performance-configuration harnesses, sharding strategies, and optimized communication patterns.
  • Debug numerical instabilities and optimize RDMA network communication using InfiniBand and RoCE.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
  • At least 5 years of experience working on performance optimization or high-performance computing systems.
  • Proficiency in CUDA or ROCm and experience with GPU profiling tools such as Nsight, nvprof, or CUPTI.
  • Familiarity with PyTorch and performance-critical model execution.
  • Experience debugging and optimizing distributed systems in multi-GPU environments.
  • Deep understanding of GPU architecture, parallel programming models, and compute kernels.
  • Master’s or PhD in Computer Science, Electrical Engineering, or a related field is preferred.
  • Experience optimizing large models for training and inference, including LLMs, VLMs, or video models, is preferred.
  • Knowledge of compiler stacks or ML compilers such as torch.compile, Triton, or XLA is preferred.
  • Open-source contributions to ML or HPC infrastructure are preferred.
  • Familiarity with cloud-scale AI infrastructure and orchestration tools such as Kubernetes is preferred.
  • Background in ML systems engineering or hardware-aware model design is preferred.

Benefits

  • Equal-opportunity and inclusive workplace
  • Collaboration with engineers and AI researchers
  • Opportunity to work on cutting-edge AI infrastructure and production generative AI systems