about 3 hours ago
Remote, Worldwide +3 moreMid Level / Senior / Staff+
Responsibilities
- Design and scale production ML systems for LLM-based applications.
- Build training and evaluation pipelines for continuous model improvement.
- Fine-tune foundation models using modern adaptation techniques.
- Optimize inference performance, latency, and GPU utilization.
- Develop high-quality training datasets and evaluation frameworks.
- Own deployment, monitoring, and reliability of ML services in production.
- Help define engineering standards while mentoring a small ML team.
Requirements
- Experience building and shipping production ML systems used by real users.
- Strong Python and PyTorch (or JAX) experience.
- Hands-on experience with LLM fine-tuning and inference optimization.
- Deep understanding of ML infrastructure, distributed training, or GPU-based systems.
- Strong engineering mindset with a focus on scalability, reliability, and code quality.
- Comfortable taking ownership in a fast-moving startup environment.
Benefits
- Founding engineering role with significant technical ownership.
- Opportunity to influence product architecture from the ground up.
- Backed by $100M while operating with the speed of a startup.
- Small, high-talent engineering team focused on solving challenging AI infrastructure problems.
- Competitive cash compensation plus equity.
- Remote-first environment with long-term growth opportunities.
