
Member of Technical Staff, Performance Optimization
Fireworks AIover 1 year ago
Base Salary
$175k - $220k/yr
Responsibilities
- Optimize system and GPU performance for high-throughput AI workloads across training and inference.
- Analyze and improve latency, throughput, memory usage, and compute efficiency.
- Profile systems to identify and resolve GPU- and kernel-level bottlenecks.
- Implement low-level optimizations using CUDA, Triton, and other performance tooling.
- Improve execution speed and resource utilization for LLM, VLM, and video-model workloads.
- Collaborate with ML researchers to co-design and tune model architectures for hardware efficiency.
- Improve support for mixed precision, quantization, and model graph optimization.
- Build and maintain performance benchmarking and monitoring infrastructure.
- Scale inference and training systems across multi-GPU and multi-node environments.
- Evaluate and integrate optimizations for emerging hardware accelerators and specialized runtimes.
- Develop asynchronous low-latency sampling, GPU kernels, distributed routing, performance-configuration harnesses, sharding strategies, and optimized communication patterns.
- Debug numerical instabilities and optimize RDMA network communication using InfiniBand and RoCE.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
- At least 5 years of experience working on performance optimization or high-performance computing systems.
- Proficiency in CUDA or ROCm and experience with GPU profiling tools such as Nsight, nvprof, or CUPTI.
- Familiarity with PyTorch and performance-critical model execution.
- Experience debugging and optimizing distributed systems in multi-GPU environments.
- Deep understanding of GPU architecture, parallel programming models, and compute kernels.
- Master’s or PhD in Computer Science, Electrical Engineering, or a related field is preferred.
- Experience optimizing large models for training and inference, including LLMs, VLMs, or video models, is preferred.
- Knowledge of compiler stacks or ML compilers such as torch.compile, Triton, or XLA is preferred.
- Open-source contributions to ML or HPC infrastructure are preferred.
- Familiarity with cloud-scale AI infrastructure and orchestration tools such as Kubernetes is preferred.
- Background in ML systems engineering or hardware-aware model design is preferred.
Benefits
- Equal-opportunity and inclusive workplace
- Collaboration with engineers and AI researchers
- Opportunity to work on cutting-edge AI infrastructure and production generative AI systems