7 hours ago
San Mateo, CA, USASenior
Responsibilities
- Design, build, and operate inference and model-serving infrastructure from development through production deployment.
- Scale AI-agent systems to operate reliably under increasing concurrency and production load.
- Identify and resolve infrastructure bottlenecks with ML and platform engineering teams.
- Optimize production systems for latency, throughput, and reliability at scale.
Requirements
- 5 or more years of experience building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.
- Hands-on experience designing and scaling inference-serving infrastructure using TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.
- Strong systems-engineering fundamentals with expertise in distributed systems, containerization, Docker, and Kubernetes.
- Experience optimizing production ML systems for latency, throughput, and reliability under high concurrency.
- Experience deploying and managing ML workloads on AWS, GCP, or Azure.
- Proficiency with monitoring, observability, and debugging tools such as Prometheus, Grafana, ELK, or distributed tracing frameworks.
- Proficiency in at least one of Python, Go, Rust, C++, or Java.
- Experience with knowledge graphs, semantic search, or graph databases such as Neo4j or Amazon Neptune is preferred.
- Familiarity with agentic AI systems, autonomous agents, or multi-step reasoning pipelines is preferred.
- Experience with enterprise data infrastructure, data pipelines, or data integration platforms is preferred.
Benefits
- The role is on-site in San Mateo, California.
