5 hours ago
San Mateo, CA, USASenior
Responsibilities
- Design, build, and own inference and model-serving infrastructure from initial architecture through production deployment.
- Scale AI-agent systems to operate reliably and efficiently under increasing concurrent load.
- Identify and resolve infrastructure bottlenecks with ML and platform engineering teams.
- Optimize production workloads for latency, throughput, and reliability.
Requirements
- 5+ years building and operating production ML inference systems, model-serving platforms, or ML infrastructure.
- Hands-on experience designing and scaling inference-serving systems with TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom solutions.
- Strong distributed systems fundamentals, including managing concurrent requests and resource allocation under load.
- Proficiency with Docker and Kubernetes for ML workloads.
- Experience deploying and managing ML systems on AWS, GCP, or Azure.
- Monitoring and observability experience with Prometheus, Grafana, ELK, or distributed tracing solutions.
- Proficiency in at least one of Python, Go, Rust, C++, or Java.
- Familiarity with knowledge graphs, semantic search, or graph databases is a plus.
- Background in agentic or autonomous AI systems, real-time inference, or enterprise data infrastructure is a plus.
Benefits
- On-site role in San Mateo, California, United States.
- Visa sponsorship is not available for this role.
