Clera

ML Infrastructure Engineer

Clera
Apply
3 days ago
San Mateo, CA, USASenior

Responsibilities

  • Own inference and model-serving infrastructure from architecture design through production deployment.
  • Build and scale reliable, efficient systems for AI agents operating under high concurrency in production.
  • Collaborate with ML and infrastructure teams on system integration and performance optimization.
  • Identify infrastructure bottlenecks and lead engineering efforts to resolve them.

Requirements

  • At least 5 years of experience building and operating production machine learning inference systems, model-serving platforms, or ML infrastructure.
  • Hands-on experience designing and scaling inference-serving infrastructure with TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.
  • Experience optimizing production ML systems for latency, throughput, and reliability at scale.
  • Strong proficiency with Docker and Kubernetes for deploying ML workloads.
  • Experience building or maintaining distributed systems that handle concurrent requests and resource allocation under load.
  • Strong command of monitoring, observability, and debugging tooling such as Prometheus, Grafana, ELK, and distributed tracing.
  • Experience deploying and managing ML systems on AWS, GCP, or Azure.
  • Proficiency in at least one of Python, Go, Rust, C++, or Java.
  • Knowledge graph, semantic search, graph database, real-time or low-latency inference, agentic AI pipeline, or enterprise data infrastructure experience is a plus.

Benefits

  • On-site work in San Mateo, California, United States.
Contact me