4 hours ago
Responsibilities
- Design, build, and operate cloud and GPU infrastructure for training and serving large models.
- Develop orchestration, job scheduling, CI/CD, developer environment, and internal tooling capabilities.
- Own reliability, observability, security, and cost optimization across compute and data systems.
- Partner with researchers and ML engineers to identify and resolve infrastructure bottlenecks.
- Establish infrastructure-as-code practices and operational standards.
Requirements
- 5 to 10+ years of experience in platform engineering, infrastructure, or site reliability engineering.
- Experience in a high-growth startup or strong engineering organization.
- Hands-on experience with AWS or GCP, Kubernetes, and infrastructure-as-code tools such as Terraform.
- Strong programming skills in Python and/or Go.
- Experience operating ML infrastructure, including GPU clusters, distributed training, and large-scale data pipelines.
- Ability to take ownership and make technical decisions in an early-stage environment.
Benefits
- Visa sponsorship is not available for this role.
- On-site position in San Francisco, California; candidates should be based in or willing to relocate to San Francisco.
