Member of Technical Staff, Inference & Serving
The Inception Company6 months ago
San Mateo, CA, USASenior
Responsibilities
- Build and optimize high-performance model-serving systems for low-latency diffusion LLM inference.
- Extend Kubernetes, Ray, and SLURM orchestration for distributed inference, evaluation, and large-batch serving.
- Implement load balancing, autoscaling, and traffic routing for model endpoints.
- Build model versioning, canary deployment, and zero-downtime rollout systems.
- Develop monitoring, alerting, and observability tooling for SLA compliance and rapid incident response.
- Collaborate with ML researchers to productionize new architectures, quantization techniques, and batching strategies.
Requirements
- BS, MS, or PhD in Computer Science, Engineering, or a related field, or equivalent experience.
- Knowledge of SGLang, vLLM, Triton Inference Server, and TensorRT-LLM.
- Systems-oriented understanding of PyTorch and TensorFlow.
- Familiarity with high-performance computing and CUDA GPU programming.
- Experience with Docker, Kubernetes, and CI/CD pipelines.
- Background in ML systems performance optimization and profiling.
- Preferred: experience with language models containing tens of billions of parameters or more.
- Preferred: experience with distributed systems, AWS, GCP, or Azure.
- Preferred: experience with Kubeflow or Airflow.
- Preferred: knowledge of quantization, distillation, speculative decoding, continuous batching, checkpointing, and resource scheduling.