5 months ago
Base Salary
$200k - $240k/yr
Responsibilities
- Build and maintain training infrastructure that supports concurrent ML model training, experiment management, and rapid iteration.
- Own the train-to-deploy handoff by exporting models to optimized inference formats, measuring accuracy and latency impact, and partnering on production deployment.
- Establish experiment tracking and lifecycle management using tools such as Weights & Biases, MLflow, or ClearML.
- Establish ML infrastructure practices on AWS, including infrastructure automation, delivery automation, observability, and cost monitoring.
- Design scalable infrastructure for applied ML and computer vision engineers and provide technical direction through architecture decisions and code.
Requirements
- At least 4 years of experience building and shipping large-scale software solutions.
- Hands-on experience building ML training pipelines in PyTorch.
- Hands-on experience with ML experiment tracking and lifecycle tools such as Weights & Biases, MLflow, or ClearML.
- Experience using AWS services such as S3, EC2, or EKS for ML workloads.
- Strong Python skills and the ability to write performant production-scale code.
- Demonstrated end-to-end ownership of infrastructure, from scoping and building through shipping and improvement.
- Strong communication skills and a bias toward shipping.
- Preferred experience with ML orchestration tools such as Ray, Sematic, Flyte, Metaflow, or Prefect.
- Preferred familiarity with GPU performance profiling and optimization tools such as Nsight or PyTorch profiler.
- Preferred background in computer vision model training.
Benefits
- Equity through Voxel’s Equity Incentive Plan.
- Total compensation includes base salary, annual bonus, and equity.
- Comprehensive health, dental, and vision insurance.
- Paid parental leave.
- Unlimited PTO and flexible work arrangements.
- Daily meals in the office, team events, and an annual company onsite.
