12 hours ago
Responsibilities
- Lead complex, ambiguous machine learning projects end to end.
- Bring foundation-model architectures from research prototypes to planetary-scale production deployment.
- Own inference efficiency, hardware/software co-design, systems architecture, and tooling.
- Build and optimize inference systems on Apple’s CloudOS and Private Cloud Compute infrastructure.
- Partner with Foundation Model Research, external partners, platform teams, and security teams.
- Build and operate high-throughput services at large distributed scale.
- Help set technical direction for surrounding engineers.
Requirements
- Proficiency in PyTorch or JAX.
- Experience working with inference frameworks.
- Experience with Python, Rust, Go, or similar programming languages.
- Proficiency deploying applications on cloud platforms such as AWS or GCP using Kubernetes and Docker.
- Experience leading complex, ambiguous machine learning projects end to end.
- Hands-on experience with LLM inference stacks is preferred.
- Working knowledge of GPU or TPU programming concepts is preferred.
- Experience building production systems in Go or Python is preferred.
- Strong knowledge of deep learning architectures, including Transformers, encoder/decoder models, and multimodal variants, is preferred.
- Experience with TensorRT-LLM, vLLM, SGLang, TGI, or Nvidia Triton Server is preferred.
- An MS in Computer Science, Machine Learning, Artificial Intelligence, Data Science, or a related field is preferred.
Categories
About Apple
We’re a diverse collective of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. And the same innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it. This is where your work can make a difference in people’s lives. Including your own. Apple is an equal opportunity employer that is committed to inclusion and diversity. Visit apple.com/careers to learn more.