20 hours ago
Responsibilities
- Evaluate Siri and AI/ML models at large scale to produce offline insights that improve model development and end-user experiences.
- Curate and evolve high-quality evaluation datasets for state-of-the-art models.
- Define and measure evaluation coverage for large language models and agentic systems.
- Address novel evaluation challenges involving personalized user experiences while upholding strict privacy standards.
- Collaborate with cross-functional teams and drive alignment on large-scale ML product or platform outcomes.
- Apply systems engineering knowledge to understand interdependencies between machine learning and software components.
Requirements
- 7+ years of professional experience applying machine learning to real-world problems and crafting scalable data solutions in natural language products.
- Proven experience managing large-scale datasets for machine learning training and/or evaluation.
- Excellent programming skills in Python.
- MS or PhD in Machine Learning, Computer Science, or equivalent experience in a related field.
- Deep knowledge of conversational AI and the end-to-end machine learning product lifecycle.
- Expertise defining and measuring evaluation coverage for large language models and agentic systems.
- Track record delivering large-scale, cross-functional machine learning product or platform outcomes.
- Strong problem-solving, critical thinking, communication, and systems engineering skills.
Tech Stack
Categories
About Apple
Apple designs and sells consumer electronics, software, and services for consumers and professionals worldwide, including iPhone, Mac, iPad, Apple Watch, and AirPods, plus platforms like iOS/macOS and services such as the App Store, iCloud, Music, and TV+. Its business combines device sales with services and subscriptions and in-house silicon design. Founded in 1976, Apple is headquartered in Cupertino, California, and trades on NASDAQ as AAPL.
