29 days ago
Palo Alto, CA, USASenior
Responsibilities
- Optimize system and GPU performance for high-throughput AI workloads across multi-node training and inference.
- Analyze and improve latency, throughput, memory usage, and compute efficiency.
- Profile systems to identify and resolve GPU- and kernel-level bottlenecks.
- Implement low-level optimizations using CUDA, Triton, and other performance tooling.
- Improve support for mixed precision, quantization, and model graph optimization.
- Build and maintain performance benchmarking and monitoring infrastructure.
- Scale inference and training systems across multi-GPU and multi-node environments.
- Own and evolve the inference engine for reliability and performance at scale.
- Develop and optimize runtime inference services for large-scale AI applications.
- Implement robust, fault-tolerant systems for data ingestion and processing.
Requirements
- 5+ years of experience writing high-quality, high-performance code.
- Familiarity with NVIDIA GPU architecture and CUDA.
- Fluency in the LLM serving stack, from kernels and quantization through schedulers and autoscaling.
- Research-leaning or systems background in LLM, Speech-to-Text, Text-to-Speech, or Speech-to-Speech inference, with demonstrable work.
- Record of shipping research or systems that other people build on in a lab or industry setting.
- Experience serving low-precision FP4 or FP8 models, using multiple LoRA adapters in one model instance, or distributing models across several GPU nodes is a plus.
- Experience developing large-scale, high-load production systems is a plus.
- Experience maintaining or contributing to open-source ML projects is a plus.
- Experience managing machine learning workloads on Kubernetes clusters is a plus.
- Experience with InfiniBand or RoCE networking, bare-metal provisioning, and lifecycle management is a plus.
- Experience operating large-scale AI training or inference clusters, monitoring hardware health, detecting predictive failures, or working with distributed storage systems is a plus.
Tech Stack
Categories
About Sanas.ai
Sanas builds real-time speech AI that translates accents and enhances clarity on live conversations, helping multilingual agents be more easily understood. It sells the technology as SaaS via APIs, call-center integrations, and an agent desktop app to contact centers, BPOs, and enterprise support and sales teams. The company is privately held and headquartered in Palo Alto, California.
