about 2 hours ago
Base Salary
$200k - $230k/yr
Responsibilities
- Architect, build, and scale the cloud infrastructure for ML training, inference, and data pipelines.
- Design resilient systems for model deployment, evaluation, and monitoring.
- Own observability across the stack, including logging, metrics, tracing, and alerting.
- Troubleshoot production issues and improve performance and efficiency.
- Collaborate with ML engineers and cross-functional teams to integrate models with data pipelines.
Requirements
- 4+ years of experience in building and scaling infrastructure in production-distributed systems.
- Strong backend software engineering fundamentals with proficiency in Python and TypeScript.
- Hands-on experience with AWS and PostgreSQL.
- Solid understanding of observability, reliability, and production incident response.
- Comfortable with ambiguity and high ownership in a startup environment.
- Interest in growing into ML infrastructure; prior ML ops/infra experience is not required.
- Nice to have: exposure to inference engines or Kubernetes.
Benefits
- Beautiful new office at 345 Hudson Street.
- Unlimited PTO.
- 100% paid employee health benefit options.
- Employer-funded 401(k) match.
- Competitive parental leave.
- Free lunch and a pantry full of snacks.
