
Member of Technical Staff, Performance Optimization
Fireworks AIover 1 year ago
Base Salary
$175k - $220k/yr
Responsibilities
- Optimize system and GPU performance for high-throughput AI training and inference workloads.
- Analyze and improve latency, throughput, memory usage, compute efficiency, execution speed, and resource utilization.
- Profile systems to identify and resolve GPU- and kernel-level bottlenecks using low-level optimizations.
- Collaborate with ML researchers to co-design and tune model architectures for hardware efficiency.
- Improve support for mixed precision, quantization, and model graph optimization.
- Build and maintain performance benchmarking and monitoring infrastructure.
- Scale inference and training systems across multi-GPU and multi-node environments.
- Evaluate and integrate optimizations for emerging hardware accelerators and specialized runtimes.
- Develop low-latency LLM sampling, GPU kernels, distributed routing, performance-configuration harnesses, model sharding, RDMA communication optimizations, and numerical-stability fixes.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
- At least 5 years of experience in performance optimization or high-performance computing systems.
- Proficiency with CUDA or ROCm and GPU profiling tools such as Nsight, nvprof, or CUPTI.
- Familiarity with PyTorch and performance-critical model execution.
- Experience debugging and optimizing distributed systems in multi-GPU environments.
- Deep understanding of GPU architecture, parallel programming models, and compute kernels.
- Master’s or PhD in a relevant field is preferred.
- Experience optimizing large models for training and inference, knowledge of compiler stacks or ML compilers, open-source ML or HPC infrastructure contributions, Kubernetes familiarity, and ML systems or hardware-aware model design experience are preferred.
Benefits
- Equal-opportunity employer committed to an inclusive environment.
Tech Stack
Categories
About Fireworks AI
Fireworks AI builds a generative AI platform for developers and enterprises to train, fine-tune, and serve open models for production use across text, image, audio, embeddings, and multimodal workloads. It offers managed inference and tooling via APIs on globally distributed infrastructure, with a usage-based SaaS model. Founded in 2022 and headquartered in San Mateo, CA, Fireworks AI is a privately held, Series D company.