1 day ago
Cupertino, CA, USA +3 moreMid Level
H1B sponsor
Responsibilities
- Design, build, and maintain automated evaluation systems for AI feature quality at scale.
- Develop evaluation architectures and methodologies for conversation- and session-level analysis rather than only single outputs.
- Create adversarial test suites and stress tests targeting multi-turn failure modes such as context degradation, goal drift, and compounding errors.
- Build evaluation frameworks, rubrics, quality assessment reports, adversarial test case libraries, and multi-turn stress-test pipelines.
- Own technical direction for evaluation efforts across multiple features or product areas.
- Evaluate conversational AI, dialogue systems, and agentic workflows, including turn-level and session-level automated scoring.
- Design tests for tool-use reliability, function-calling accuracy, and agent planning quality.
- Communicate evaluation findings and model-readiness assessments to cross-functional partners.
Requirements
- Bachelor’s degree in Computer Science, Machine Learning, Statistics, or a related field.
- At least 4 years of experience building or significantly extending ML evaluation systems, including evaluation benchmarks or quality assessment frameworks for sequential or multi-step AI outputs.
- Experience independently defining evaluation architecture and methodology for AI or ML systems using conversations or sessions as the unit of analysis.
- Experience designing adversarial or red-teaming methodologies for ML models or AI-powered features, including multi-turn failure scenarios.
- Production or near-production experience with Python and ML frameworks such as PyTorch or TensorFlow.
- Track record of owning technical direction across multiple features or product areas.
- Preferred experience evaluating user-facing AI features in consumer applications and connecting technical metrics to user-perceived quality.
- Preferred familiarity with productivity software or creative tools and workflow-based output assessment.
- Preferred experience aligning automated and human evaluation methods, including inter-annotator agreement analysis and bias detection.
- Preferred experience designing evaluation systems that scale across multiple features without bespoke solutions.
- Preferred experience evaluating API-based and custom-trained AI models.
- Preferred experience automating evaluation data generation and analysis.
- Preferred familiarity with LangChain, LangGraph, CrewAI, AutoGen, LangSmith, Braintrust, and Arize.
- Graduate degree in a relevant field is preferred.
Categories
About Apple
Apple designs and sells consumer electronics, software, and services for consumers and professionals worldwide, including iPhone, Mac, iPad, Apple Watch, and AirPods, plus platforms like iOS/macOS and services such as the App Store, iCloud, Music, and TV+. Its business combines device sales with services and subscriptions and in-house silicon design. Founded in 1976, Apple is headquartered in Cupertino, California, and trades on NASDAQ as AAPL.
