5 hours ago
Base Salary
$170k - $230k/yr
Responsibilities
- Own and maintain CI/CD infrastructure for ML models and inference services.
- Build automated testing, model validation, deployment verification, rollback, and monitoring systems.
- Accelerate CI execution through parallelization, caching, test selection, and efficient compute usage.
- Develop automated model quality and performance benchmarks across models, GPU architectures, and configurations.
- Automate validation of pricing, billing configurations, API schemas, and deployments.
- Build AI-powered and agentic engineering workflows that diagnose CI failures, identify regressions, and propose fixes.
- Identify and automate repetitive ML engineering tasks to reduce developer toil.
Requirements
- 5+ years of software engineering experience with proficiency in Python and production infrastructure or developer tooling.
- Experience designing and operating CI/CD systems using GitHub Actions or comparable technologies.
- Strong understanding of automated testing, build systems, dependency management, caching, and parallel execution.
- Experience with Docker, containerized workloads, and cloud infrastructure.
- Ability to design reliable distributed systems and debug complex infrastructure failures.
- Strong understanding of observability, including logs, metrics, tracing, and automated alerting.
- Experience with ML infrastructure, PyTorch, GPU workloads, or model-serving systems.
- Familiarity with NVIDIA GPU architectures and multi-GPU environments.
- Experience building inference benchmarks or ML quality evaluation frameworks.
- Experience with AI coding agents such as Codex or Claude Code, or custom AI engineering agents.
- Experience optimizing CI/CD at scale, including distributed test execution and ephemeral compute environments.
- Experience developing internal developer platforms or infrastructure-as-code tooling.
Benefits
- Health, dental, and vision insurance for US employees.
- Learning and growth opportunities.
- Regular team events and offsites.
- Interesting and challenging work.
Categories
About fal
Fal builds a generative media platform that gives developers a unified API to run state-of-the-art image, video, and audio models. It provides serverless GPUs, high-performance inference, and dedicated compute clusters so teams can customize, deploy, and scale models in production. The company is privately held and headquartered in San Francisco, serving both startups and enterprises through a commercial API and managed infrastructure.
