3 days ago
San Mateo, CA, USASenior
Responsibilities
- Own inference and model-serving infrastructure from architecture design through production deployment.
- Build and scale reliable, efficient systems for AI agents operating under high concurrency in production.
- Collaborate with ML and infrastructure teams on system integration and performance optimization.
- Identify infrastructure bottlenecks and lead engineering efforts to resolve them.
Requirements
- At least 5 years of experience building and operating production machine learning inference systems, model-serving platforms, or ML infrastructure.
- Hands-on experience designing and scaling inference-serving infrastructure with TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.
- Experience optimizing production ML systems for latency, throughput, and reliability at scale.
- Strong proficiency with Docker and Kubernetes for deploying ML workloads.
- Experience building or maintaining distributed systems that handle concurrent requests and resource allocation under load.
- Strong command of monitoring, observability, and debugging tooling such as Prometheus, Grafana, ELK, and distributed tracing.
- Experience deploying and managing ML systems on AWS, GCP, or Azure.
- Proficiency in at least one of Python, Go, Rust, C++, or Java.
- Knowledge graph, semantic search, graph database, real-time or low-latency inference, agentic AI pipeline, or enterprise data infrastructure experience is a plus.
Benefits
- On-site work in San Mateo, California, United States.
