Clera

ML Infrastructure Engineer

Clera
Apply
5 hours ago
San Mateo, CA, USASenior

Responsibilities

  • Design, build, and own inference and model-serving infrastructure from initial architecture through production deployment.
  • Scale AI-agent systems to operate reliably and efficiently under increasing concurrent load.
  • Identify and resolve infrastructure bottlenecks with ML and platform engineering teams.
  • Optimize production workloads for latency, throughput, and reliability.

Requirements

  • 5+ years building and operating production ML inference systems, model-serving platforms, or ML infrastructure.
  • Hands-on experience designing and scaling inference-serving systems with TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom solutions.
  • Strong distributed systems fundamentals, including managing concurrent requests and resource allocation under load.
  • Proficiency with Docker and Kubernetes for ML workloads.
  • Experience deploying and managing ML systems on AWS, GCP, or Azure.
  • Monitoring and observability experience with Prometheus, Grafana, ELK, or distributed tracing solutions.
  • Proficiency in at least one of Python, Go, Rust, C++, or Java.
  • Familiarity with knowledge graphs, semantic search, or graph databases is a plus.
  • Background in agentic or autonomous AI systems, real-time inference, or enterprise data infrastructure is a plus.

Benefits

  • On-site role in San Mateo, California, United States.
  • Visa sponsorship is not available for this role.
Contact me