10 months ago
Base Salary
$165k - $330k/yr
Responsibilities
- Build and maintain scalable infrastructure for deploying and operating machine learning models
- Establish infrastructure standards and best practices for reliability and performance
- Automate processes, particularly for managing CI/CD pipelines
- Own products and projects end to end, including project specification, execution, and user-focused decision-making
- Collaborate with cross-functional teams to translate project requirements into technical solutions
- Mentor junior team members and contribute to organizational knowledge sharing
- Manage tradeoffs and select appropriate tools while avoiding unnecessary complexity
- Work on initiatives involving multi-cloud capacity management, GPU inference, multi-node inference, and fractional GPU model serving
Requirements
- Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field
- Extensive experience with Kubernetes
- Experience building and maintaining scalable infrastructure
- Experience with infrastructure-as-code tools such as Terraform, CloudFormation, or Pulumi
- Experience with CI/CD tooling such as GitHub Actions, GitLab CI, Circle CI, or Jenkins
- Relevant open-source observability experience with tools such as Prometheus, the ELK stack, Grafana, or OpenTelemetry is a plus
- Ability to own projects end to end, from project specification through execution
- Prior machine learning experience is not required, but candidates should be open to learning about it
Benefits
- 100% coverage of medical, dental, and vision insurance for employees and dependents
- Flexible PTO, including a company-wide Winter Break with offices closed from Christmas Eve through New Year's Day
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Exposure to a variety of machine learning startups and networking opportunities
Tech Stack
Categories
About Baseten
Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.
