15 hours ago
Responsibilities
- Optimize machine learning model training and inference across the ML stack.
- Improve model performance, training speed, and inference speed on Apple Silicon.
- Write production-level code to train and deploy models.
- Work across data, modeling, evaluation, deployment, ML infrastructure, inference, and framework teams.
- Optimize and deploy LLM models, including techniques such as quantization, KV Cache, and Speculative Decoding.
Requirements
- Experience with the model lifecycle, including training, evaluation, and deployment.
- Strong understanding of machine learning model architectures such as Transformers and CNNs and of ML training loops.
- Strong proficiency in Python and an ML framework such as PyTorch.
- Bachelor's degree in Computer Science, Engineering, or a related discipline, or equivalent industry or project experience.
- Experience with agentic AI-assisted coding.
- Preferred expertise in ML and LLM optimization, including quantization, KV Cache, and Speculative Decoding.
- Preferred familiarity with FSDP, DDP, and other parallelism methods.
- Preferred experience with an LLM training or evaluation library such as HuggingFace Transformers, lm evaluation harness, or Megatron-LM.
- Preferred proficiency in a compiled programming language such as Swift, C, C++, or Java.
- Preferred experience collaborating on large inter-team projects.
Categories
About Apple
Apple designs and sells consumer electronics, software, and services for consumers and professionals worldwide, including iPhone, Mac, iPad, Apple Watch, and AirPods, plus platforms like iOS/macOS and services such as the App Store, iCloud, Music, and TV+. Its business combines device sales with services and subscriptions and in-house silicon design. Founded in 1976, Apple is headquartered in Cupertino, California, and trades on NASDAQ as AAPL.
