The Inception Company

Member of Technical Staff, Inference & Serving

The Inception Company
Apply
6 months ago
San Mateo, CA, USASenior

Responsibilities

  • Build and optimize high-performance model-serving systems for low-latency diffusion LLM inference.
  • Extend Kubernetes, Ray, and SLURM orchestration for distributed inference, evaluation, and large-batch serving.
  • Implement load balancing, autoscaling, and traffic routing for model endpoints.
  • Build model versioning, canary deployment, and zero-downtime rollout systems.
  • Develop monitoring, alerting, and observability tooling for SLA compliance and rapid incident response.
  • Collaborate with ML researchers to productionize new architectures, quantization techniques, and batching strategies.

Requirements

  • BS, MS, or PhD in Computer Science, Engineering, or a related field, or equivalent experience.
  • Knowledge of SGLang, vLLM, Triton Inference Server, and TensorRT-LLM.
  • Systems-oriented understanding of PyTorch and TensorFlow.
  • Familiarity with high-performance computing and CUDA GPU programming.
  • Experience with Docker, Kubernetes, and CI/CD pipelines.
  • Background in ML systems performance optimization and profiling.
  • Preferred: experience with language models containing tens of billions of parameters or more.
  • Preferred: experience with distributed systems, AWS, GCP, or Azure.
  • Preferred: experience with Kubeflow or Airflow.
  • Preferred: knowledge of quantization, distillation, speculative decoding, continuous batching, checkpointing, and resource scheduling.
The Inception Company

About The Inception Company

51-200 employees
Contact me