Apple

ML Engineer - Automated Evaluation and Adversarial Design

Apple
Apply
1 day ago
Cupertino, CA, USA +3 moreMid Level
H1B sponsor

Responsibilities

  • Design, build, and maintain automated evaluation systems for AI feature quality at scale.
  • Develop evaluation architectures and methodologies for conversation- and session-level analysis rather than only single outputs.
  • Create adversarial test suites and stress tests targeting multi-turn failure modes such as context degradation, goal drift, and compounding errors.
  • Build evaluation frameworks, rubrics, quality assessment reports, adversarial test case libraries, and multi-turn stress-test pipelines.
  • Own technical direction for evaluation efforts across multiple features or product areas.
  • Evaluate conversational AI, dialogue systems, and agentic workflows, including turn-level and session-level automated scoring.
  • Design tests for tool-use reliability, function-calling accuracy, and agent planning quality.
  • Communicate evaluation findings and model-readiness assessments to cross-functional partners.

Requirements

  • Bachelor’s degree in Computer Science, Machine Learning, Statistics, or a related field.
  • At least 4 years of experience building or significantly extending ML evaluation systems, including evaluation benchmarks or quality assessment frameworks for sequential or multi-step AI outputs.
  • Experience independently defining evaluation architecture and methodology for AI or ML systems using conversations or sessions as the unit of analysis.
  • Experience designing adversarial or red-teaming methodologies for ML models or AI-powered features, including multi-turn failure scenarios.
  • Production or near-production experience with Python and ML frameworks such as PyTorch or TensorFlow.
  • Track record of owning technical direction across multiple features or product areas.
  • Preferred experience evaluating user-facing AI features in consumer applications and connecting technical metrics to user-perceived quality.
  • Preferred familiarity with productivity software or creative tools and workflow-based output assessment.
  • Preferred experience aligning automated and human evaluation methods, including inter-annotator agreement analysis and bias detection.
  • Preferred experience designing evaluation systems that scale across multiple features without bespoke solutions.
  • Preferred experience evaluating API-based and custom-trained AI models.
  • Preferred experience automating evaluation data generation and analysis.
  • Preferred familiarity with LangChain, LangGraph, CrewAI, AutoGen, LangSmith, Braintrust, and Arize.
  • Graduate degree in a relevant field is preferred.

Tech Stack

PythonPyTorchTensorFlow
Apple

About Apple

10,000+ employees

Apple designs and sells consumer electronics, software, and services for consumers and professionals worldwide, including iPhone, Mac, iPad, Apple Watch, and AirPods, plus platforms like iOS/macOS and services such as the App Store, iCloud, Music, and TV+. Its business combines device sales with services and subscriptions and in-house silicon design. Founded in 1976, Apple is headquartered in Cupertino, California, and trades on NASDAQ as AAPL.

Contact me