12 hours ago
San Mateo, CA, USASenior
Responsibilities
- Design, build, and scale inference and model-serving infrastructure through production deployment.
- Optimize ML infrastructure for latency, throughput, and reliability under high concurrency.
- Collaborate with ML and infrastructure teams to integrate systems and identify performance bottlenecks.
- Drive infrastructure solutions across a fast-moving, cross-functional team.
Requirements
- At least 5 years of experience building and operating ML inference systems, model-serving platforms, or ML infrastructure in production.
- Hands-on experience with TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom serving solutions.
- Strong distributed systems fundamentals, including Docker containerization and Kubernetes orchestration.
- Experience with production monitoring and observability tools such as Prometheus and Grafana.
- Experience deploying and managing ML workloads on AWS, GCP, or Azure.
- Proficiency in at least one of Python, Go, Rust, C++, or Java.
- Preferred experience with knowledge graphs, semantic search, graph databases, real-time or low-latency inference, agentic or multi-step AI pipelines, or enterprise data integration and pipeline infrastructure.
Benefits
- On-site role in San Mateo, California, United States.
