3 months ago
Remote, India +2 moreMid Level
Responsibilities
- Own availability, latency, and throughput SLOs for production generative media model APIs.
- Build monitoring, alerting, and observability for ML-specific failures, output degradation, pipeline breakage, and model regressions.
- Harden model deployment with canary releases, shadow testing, automated rollbacks, and validation gates.
- Improve model-fleet security through secure serving, abuse and misuse detection, rate limiting, and adversarial-use protection.
- Operationalize content moderation pipelines, safety classifiers, and inference-time guardrails.
- Lead incident response and postmortems for model API outages and degradations, and prevent recurrence.
- Improve capacity planning, autoscaling, and GPU fleet efficiency for variable inference workloads.
- Partner with model and infrastructure teams to incorporate reliability, security, and safety requirements into model onboarding.
Requirements
- At least 3 years of professional experience, including 1 year operating production ML or high-scale API systems, ideally with on-call ownership.
- Strong fundamentals in distributed systems, networking, observability, and incident management.
- Working knowledge of diffusion models, transformers, and their production failure modes.
- Familiarity with ML security and safety practices, abuse prevention, content safety, or trust and safety engineering is a strong plus.
- Ability to automate, measure system behavior, and conduct blameless postmortems.
Benefits
- Hybrid/remote role based in India, Australia, or New Zealand.
- Access to a large GPU cluster for inference and evaluation.
Tech Stack
Categories
ML EngineeringSite Reliability
About fal
Fal builds a generative media platform that gives developers a unified API to run state-of-the-art image, video, and audio models. It provides serverless GPUs, high-performance inference, and dedicated compute clusters so teams can customize, deploy, and scale models in production. The company is privately held and headquartered in San Francisco, serving both startups and enterprises through a commercial API and managed infrastructure.
