1 day ago
Base Salary
$388k - $619k/yr
Responsibilities
- Build observability frameworks and platform capabilities for metrics, logs, and distributed traces across online inference, batch scoring, feature pipelines, and agent orchestration.
- Develop reusable primitives for monitoring model performance, data quality, drift, and degradation.
- Build evaluation frameworks for LLM and agentic systems covering response quality, grounding, hallucination, task success, tool use, trajectory correctness, LLM-as-a-judge, and human-in-the-loop scoring.
- Lead build-versus-buy evaluations and own SDKs, connectors, and APIs for vendor tooling integrations.
- Create reusable libraries, SDKs, and templates that make observability and evaluation the default for new systems.
- Provide dashboarding, alerting, and SLO/SLI building blocks for model performance, latency, cost, and reliability.
- Partner with engineering, product, machine learning, and data teams to turn needs into reusable platform capabilities.
Requirements
- Experience in software, AI/ML, or platform engineering, including production observability, monitoring, or ML/LLM evaluation.
- A proven track record designing standards, libraries, SDKs, or frameworks adopted by other teams.
- Strong coding skills in Python and at least one of Java, Go, or Scala, with experience building production services.
- Practical experience operating ML models in production through online serving and/or batch processing, including drift, model performance, and data quality.
- Hands-on experience with modern observability stacks such as Prometheus/Grafana, Datadog, OpenTelemetry, ELK/OpenSearch, or Jaeger/Tempo.
- Understanding of distributed systems, microservices, and at least one major cloud platform: AWS, GCP, or Azure.
- Ability to work cross-functionally with ML, data, infrastructure, and product teams and communicate clearly about system behavior, quality, and risk.
- Experience with vendor integration and VPC deployment.
- Experience using AI tools as a core part of design, development, testing, and code review workflows.
- Preferred: hands-on experience with ML/LLM observability and evaluation tools such as Arize, Braintrust, LangFuse, Weights & Biases, Galileo, Vertex AI Model Monitoring, or SageMaker Model Monitor.
- Preferred: experience building or shipping LLM/GenAI applications and evaluating prompt/result logging, evaluation metrics, LLM-as-a-judge, and human-in-the-loop review.
- Preferred: experience evaluating agentic systems, including tool use, multi-step reasoning, and trajectory or task-success measurement.
Benefits
- Annual salary range of $388,000.00-$619,000.00, with compensation selectable between salary and stock options and no bonuses.
- Health plans, mental health support, a 401(k) retirement plan with employer match, stock option program, disability programs, health savings and flexible spending accounts, family-forming benefits, and life and serious injury benefits.
- Paid leave programs; full-time salaried employees are immediately entitled to flexible time off.
- Equal-opportunity workplace with accommodation support available during the interview process.
Categories
About Netflix
Netflix builds and operates a global streaming service for TV series, films, games, and live programming, and produces original content through its in-house studio. It serves consumers in 190+ countries via subscription and ad-supported plans on internet-connected devices. Founded in 1997 and headquartered in Los Gatos, California, Netflix is a public company traded on NASDAQ (NFLX).
