Member of Technical Staff, Backend, LLM Applications
The Inception Company6 months ago
San Mateo, CA, USASenior
Responsibilities
- Design, build, and operate scalable backend services and model-serving infrastructure for diffusion LLMs.
- Implement load balancing, autoscaling, and traffic routing for model endpoints.
- Build model versioning, canary deployment, and zero-downtime rollout systems.
- Develop monitoring, alerting, and observability tooling for SLA compliance and incident response.
- Benchmark serving frameworks and hardware configurations to inform infrastructure decisions.
- Manage cloud infrastructure, GPU instances, infrastructure as code, and deployment automation.
Requirements
- Bachelor's, master's, or doctoral degree in Computer Science or a related field, or equivalent experience.
- At least 5 years of experience building production backend systems.
- Strong proficiency in Python, including asynchronous programming and concurrent systems.
- Strong understanding of distributed systems, networking, and load balancing at scale.
- Familiarity with Kubernetes, CI/CD pipelines, and cloud infrastructure using AWS and/or Azure.
- Preferred experience serving LLMs or other large generative models in production at scale.
- Preferred experience with GPU instance management, cloud cost optimization, Terraform, Prometheus, Grafana, vLLM, Triton Inference Server, or TensorRT-LLM.