
Member of Technical Staff, Cloud Infrastructure
Fireworks AIabout 1 year ago
Base Salary
$175k - $220k/yr
Responsibilities
- Architect and build scalable, resilient, high-performance backend infrastructure for distributed training, inference, and data-processing pipelines.
- Design and implement backend services including job schedulers, resource managers, autoscalers, and model-serving layers.
- Lead technical design discussions, mentor engineers, and establish best practices for large-scale ML infrastructure.
- Optimize compute costs, storage lifecycle management, network performance, system reliability, and operational efficiency.
- Collaborate with ML, DevOps, product, and infrastructure teams to translate research and product needs into robust solutions.
- Evaluate and integrate cloud-native and open-source technologies to improve platform capabilities and reliability.
- Own systems end to end from design through deployment and observability, with emphasis on availability, scalability, fault tolerance, disaster recovery, and performance.
- Develop and maintain monitoring, alerting, logging, and tracing solutions for system health and performance.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- At least 5 years of experience designing and building backend infrastructure in cloud environments.
- Proven experience with ML infrastructure and tooling such as PyTorch, TensorFlow, Vertex AI, SageMaker, or Kubernetes.
- Strong software development skills in Python or C++.
- Deep understanding of distributed systems fundamentals, including scheduling, orchestration, storage, networking, and compute optimization.
- Master’s or PhD in Computer Science or a related field is preferred.
- Experience leading infrastructure projects for large-scale ML/AI workloads or high-throughput systems is preferred.
- Familiarity with infrastructure-as-code and CI/CD tooling such as Terraform, ArgoCD, or GitOps is preferred.
- A track record of improving system performance, reliability, and cost efficiency is preferred; open-source cloud or ML infrastructure contributions are a plus.
Benefits
- Work on cutting-edge AI infrastructure and large-scale generative AI systems.
- Collaborate with engineers and AI researchers in a fast-growing, inclusive environment.
- The role involves ownership and direct impact with no bureaucracy.
- Equal-opportunity employer committed to diversity and inclusion.