6 months ago
San Francisco, CA, USA or New York, NY, USASenior
Base Salary
$200k - $400k/yr
Responsibilities
- Build tools and infrastructure for understanding and intentionally designing models at industry scale.
- Extend and support training infrastructure for large training runs.
- Translate cutting-edge interpretability research into production-ready tools.
- Optimize pipelines and infrastructure for frontier model interpretability, training, and inference.
- Integrate new machine learning workflows and pipelines into the product and deploy them to customers.
- Ensure system reliability, reproducibility, and performance.
Requirements
- At least five years of experience in ML infrastructure, research engineering, or systems programming.
- Expertise in Python, PyTorch or JAX, and distributed systems.
- Experience deploying and maintaining machine learning systems at scale.
- Ability to work across research and engineering boundaries.
- Commitment to understanding how models work internally and using that understanding to improve their reliability and usefulness.
- Open-source ML infrastructure contributions are preferred.
- Startup or frontier lab experience in fast-moving teams is preferred.
Benefits
- Market-competitive salary, equity, and competitive benefits.
- In-person work five days per week in either the San Francisco HQ or New York office, with one company-wide remote week per month.