GrepJob
fal

Machine Learning Engineer, Reliability

fal
Apply
about 2 months ago
Remote, WorldwideMid Level
H1B Sponsor

Responsibilities

  • Own availability, latency, and throughput SLOs for a large fleet of production generative media model APIs.
  • Build monitoring, alerting, and observability for ML-specific failures, output degradation, pipeline breakage, and model regressions.
  • Harden model deployment with canary releases, shadow testing, automated rollbacks, and validation gates.
  • Drive secure model serving, abuse and misuse detection, rate limiting, and protection against adversarial usage.
  • Operationalize content moderation pipelines, safety classifiers, and inference-time guardrails.
  • Lead incident response and postmortems for model API outages and degradations and prevent recurrence.
  • Improve capacity planning, autoscaling, and GPU fleet efficiency for variable inference workloads.
  • Partner with model and infrastructure teams to incorporate reliability, security, and safety requirements into model onboarding.

Requirements

  • At least three years of professional experience, including at least one year operating production ML or high-scale API systems.
  • Strong systems fundamentals in distributed systems, networking, observability, and incident management.
  • Working knowledge of generative models such as diffusion models and transformers and their production failure modes.
  • Familiarity with ML security and safety practices, with abuse prevention, content safety, or trust and safety engineering experience valued as a strong plus.
  • Experience with on-call ownership is ideal.

Benefits

  • Remote work based in India, Australia, or New Zealand.
  • Access to a large GPU cluster for inference and evaluation.