The Inception Company

Member of Technical Staff, Backend, LLM Applications

The Inception Company
Apply
6 months ago
San Mateo, CA, USASenior

Responsibilities

  • Design, build, and operate scalable backend services and model-serving infrastructure for diffusion LLMs.
  • Implement load balancing, autoscaling, and traffic routing for model endpoints.
  • Build model versioning, canary deployment, and zero-downtime rollout systems.
  • Develop monitoring, alerting, and observability tooling for SLA compliance and incident response.
  • Benchmark serving frameworks and hardware configurations to inform infrastructure decisions.
  • Manage cloud infrastructure, GPU instances, infrastructure as code, and deployment automation.

Requirements

  • Bachelor's, master's, or doctoral degree in Computer Science or a related field, or equivalent experience.
  • At least 5 years of experience building production backend systems.
  • Strong proficiency in Python, including asynchronous programming and concurrent systems.
  • Strong understanding of distributed systems, networking, and load balancing at scale.
  • Familiarity with Kubernetes, CI/CD pipelines, and cloud infrastructure using AWS and/or Azure.
  • Preferred experience serving LLMs or other large generative models in production at scale.
  • Preferred experience with GPU instance management, cloud cost optimization, Terraform, Prometheus, Grafana, vLLM, Triton Inference Server, or TensorRT-LLM.

Tech Stack

The Inception Company

About The Inception Company

51-200 employees
Contact me