over 1 year ago
Base Salary
$350k - $600k/yr
Responsibilities
- Lead the design and development of the core ML stack for large-scale AI systems.
- Build high-concurrency container-based evaluations and deterministic sandbox execution with checkpoint restoration.
- Develop performant, flexible inference systems supporting model introspection, intervention, steering, configurable sampling, and 400B+ parameter models.
- Build distributed reinforcement-learning training and rollout systems supporting thousands of concurrent rollouts across machines.
- Set code culture, tooling, infrastructure direction, and engineering best practices across the organization.
- Build internal tools and advise team members on infrastructure challenges.
Requirements
- Exceptional programming ability and fluency in Python.
- Strong knowledge of GPUs and other accelerators, including low-level performance optimization and parallel programming.
- Experience engineering at scale with distributed systems, reliability, and architecture design.
- Demonstrated leadership in code quality, system primitives, complexity management, and scalability.
- Experience with LLM pipelines involving multiple specialized LLMs is a bonus.
- Experience with open-source community management is a bonus.
Benefits
- Benefits are provided in addition to the stated salary.
- The role is based in San Francisco and the team is enthusiastic about working together in person.
- The organization is open to sponsoring international visas.
Tech Stack
Categories
About Transluce
Transluce is an independent nonprofit research lab in San Francisco, founded in 2024, that builds open-source tools to analyze and evaluate advanced AI systems. It develops automated assessments of behaviors like honesty, misreporting, and evaluation awareness on open-weight models, then works with frontier AI labs and governments to apply vetted procedures. The lab focuses on public-interest AI transparency, reliability, and scalable oversight.
