about 2 months ago
Remote, WorldwideMid Level
H1B Sponsor
Responsibilities
- Own availability, latency, and throughput SLOs for a large fleet of production generative media model APIs.
- Build monitoring, alerting, and observability for ML-specific failures, output degradation, pipeline breakage, and model regressions.
- Harden model deployment with canary releases, shadow testing, automated rollbacks, and validation gates.
- Drive secure model serving, abuse and misuse detection, rate limiting, and protection against adversarial usage.
- Operationalize content moderation pipelines, safety classifiers, and inference-time guardrails.
- Lead incident response and postmortems for model API outages and degradations and prevent recurrence.
- Improve capacity planning, autoscaling, and GPU fleet efficiency for variable inference workloads.
- Partner with model and infrastructure teams to incorporate reliability, security, and safety requirements into model onboarding.
Requirements
- At least three years of professional experience, including at least one year operating production ML or high-scale API systems.
- Strong systems fundamentals in distributed systems, networking, observability, and incident management.
- Working knowledge of generative models such as diffusion models and transformers and their production failure modes.
- Familiarity with ML security and safety practices, with abuse prevention, content safety, or trust and safety engineering experience valued as a strong plus.
- Experience with on-call ownership is ideal.
Benefits
- Remote work based in India, Australia, or New Zealand.
- Access to a large GPU cluster for inference and evaluation.
