7 hours ago
Responsibilities
- Design and develop automated evaluation frameworks and pipelines for AI-powered products.
- Define quality metrics and evaluation methodologies covering accuracy, relevance, groundedness, completeness, consistency, instruction following, and task completion.
- Build golden datasets, benchmark sets, regression suites, adversarial scenarios, and production-derived test sets.
- Develop Auto Eval and model-based evaluation capabilities, including LLM-as-a-Judge, calibration, and validation methods.
- Create Human-in-the-Loop evaluation processes, rubrics, annotation guidelines, grading criteria, and quality standards.
- Evaluate retrieval, context construction, prompts, model responses, tool use, APIs, and end-to-end AI experiences.
- Develop evaluation methods for conversations, personalization, recommendations, tool use, reasoning, and agentic task execution.
- Perform error analysis and failure-mode investigation to identify model, prompt, retrieval, dataset, and product improvements.
- Build reusable evaluation infrastructure, APIs, dashboards, and developer tooling for multiple AI products and teams.
- Integrate evaluation into CI/CD workflows with automated regression detection, quality gates, and release-readiness assessments.
- Connect offline evaluation results with production signals to improve evaluation coverage and product quality.
- Partner with Machine Learning, Software Engineering, Product, Quality Engineering, Human Interface, and Data Science teams across the AI product lifecycle.
Requirements
- At least 7 years of related experience in Machine Learning Engineering, ML Evaluation, Software Engineering, Data Science, Quality Engineering, or a related technical field.
- Strong Python programming skills and experience developing production-quality software, ML systems, data pipelines, or evaluation infrastructure.
- Experience developing or evaluating LLMs, Generative AI, Conversational AI, NLP, recommendation systems, or other machine-learning-driven products.
- Experience designing automated ML evaluation frameworks, metrics, benchmarks, datasets, or experimentation methodologies.
- Understanding of LLM application architectures, including prompting, embeddings, retrieval-augmented generation, tool use, and agentic workflows.
- Experience with model-based evaluation and understanding of LLM-as-a-Judge strengths and limitations.
- Experience with Human-in-the-Loop evaluation, annotation, or data-quality workflows.
- Strong understanding of statistical analysis, experimentation, sampling, and measurement methodologies.
- Experience with model error analysis, failure analysis, and root-cause investigation.
- Bachelor’s degree in Computer Science, Machine Learning, Artificial Intelligence, Data Science, Statistics, Electrical Engineering, or a related technical field, or equivalent industry experience.
- Preferred experience with production-scale LLM or Generative AI evaluation infrastructure, RAG, conversational systems, AI agents, personalization, recommendations, or multimodal AI.
- Preferred experience with golden datasets, regression suites, automated quality gates, continuous evaluation pipelines, CI/CD integration, experiment tracking, AI observability, multilingual evaluation, responsible AI, internal ML platforms, developer tooling, and distributed ML or data-processing infrastructure.
- A master’s degree in a related technical field or equivalent industry experience is preferred.
Tech Stack
Categories
About Apple
Apple designs and sells consumer electronics, software, and services for consumers and professionals worldwide, including iPhone, Mac, iPad, Apple Watch, and AirPods, plus platforms like iOS/macOS and services such as the App Store, iCloud, Music, and TV+. Its business combines device sales with services and subscriptions and in-house silicon design. Founded in 1976, Apple is headquartered in Cupertino, California, and trades on NASDAQ as AAPL.
