GrepJob
Fireworks AI

Member of Technical Staff, Cloud Infrastructure

Fireworks AI
Apply
about 1 year ago
San Mateo, CA, USA or New York, NY, USASenior
H1B Sponsor

Base Salary

$175k - $220k/yr

Responsibilities

  • Architect and build scalable, resilient, high-performance backend infrastructure for distributed training, inference, and data-processing pipelines.
  • Design and implement backend services including job schedulers, resource managers, autoscalers, and model-serving layers.
  • Lead technical design discussions, mentor engineers, and establish best practices for large-scale ML infrastructure.
  • Optimize compute costs, storage lifecycle management, network performance, system reliability, and operational efficiency.
  • Collaborate with ML, DevOps, product, and infrastructure teams to translate research and product needs into robust solutions.
  • Evaluate and integrate cloud-native and open-source technologies to improve platform capabilities and reliability.
  • Own systems end to end from design through deployment and observability, with emphasis on availability, scalability, fault tolerance, disaster recovery, and performance.
  • Develop and maintain monitoring, alerting, logging, and tracing solutions for system health and performance.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • At least 5 years of experience designing and building backend infrastructure in cloud environments.
  • Proven experience with ML infrastructure and tooling such as PyTorch, TensorFlow, Vertex AI, SageMaker, or Kubernetes.
  • Strong software development skills in Python or C++.
  • Deep understanding of distributed systems fundamentals, including scheduling, orchestration, storage, networking, and compute optimization.
  • Master’s or PhD in Computer Science or a related field is preferred.
  • Experience leading infrastructure projects for large-scale ML/AI workloads or high-throughput systems is preferred.
  • Familiarity with infrastructure-as-code and CI/CD tooling such as Terraform, ArgoCD, or GitOps is preferred.
  • A track record of improving system performance, reliability, and cost efficiency is preferred; open-source cloud or ML infrastructure contributions are a plus.

Benefits

  • Work on cutting-edge AI infrastructure and large-scale generative AI systems.
  • Collaborate with engineers and AI researchers in a fast-growing, inclusive environment.
  • The role involves ownership and direct impact with no bureaucracy.
  • Equal-opportunity employer committed to diversity and inclusion.