12 hours ago
Responsibilities
- Own major existing or new systems in Apple’s AI evaluation platform.
- Build LLM-as-judge autograders, tools that validate them against human judgment, and systems that run evaluations at scale.
- Set the direction, quality standards, and adoption strategy for owned systems.
- Build and maintain technical integrations across multiple platforms and teams.
- Interpret evaluation results, including human agreement and run-to-run consistency.
- Balance scientific research processes with product development and release lifecycles.
- Turn ambiguous product and technical requirements into focused plans.
Requirements
- BS, MS, or PhD in Computer Science, Machine Learning, or a related field.
- 8+ years of software engineering experience, including ownership of production systems used by other teams.
- Exceptional Python skills and strong understanding of system design, API design, system testing and validation, debugging, and monitoring.
- Experience shipping production code with AI coding agents while maintaining high quality.
- Preferred: experience building agentic systems, LLM-based applications, LLM-as-judge systems, and offline evaluations.
- Preferred: ability to interpret evaluation results and account for data retention and privacy constraints.
- Preferred: expertise building and maintaining technical integrations across multiple platforms in a large organization.
- Preferred: product-minded approach to ambiguous requirements and focused planning.
Tech Stack
Categories
About Apple
Apple designs and sells consumer electronics, software, and services for consumers and professionals worldwide, including iPhone, Mac, iPad, Apple Watch, and AirPods, plus platforms like iOS/macOS and services such as the App Store, iCloud, Music, and TV+. Its business combines device sales with services and subscriptions and in-house silicon design. Founded in 1976, Apple is headquartered in Cupertino, California, and trades on NASDAQ as AAPL.
