Cloud Platform Engineer
SambaNova SystemsResponsibilities
- Own the availability, latency, performance, efficiency, change management, monitoring, emergency response, and capacity planning of the production inferencing service across multiple regions.
- Participate in a balanced primary/secondary on-call rotation supporting 24/7 service reliability.
- Lead incident response, blameless post-mortems, and corrective actions for inferencing-service incidents.
- Develop monitoring, alerting, and dashboards for service health, model performance, and accelerator utilization.
- Identify performance bottlenecks and implement cost-effective auto-scaling policies.
- Manage cloud and on-premises infrastructure across AWS, GCP, and/or Azure using Infrastructure as Code.
- Build and improve CI/CD pipelines for model versions and service updates while automating operational toil.
- Forecast infrastructure needs, manage cloud costs, and optimize spending with finance and engineering teams.
- Define, measure, and report Service Level Objectives and Service Level Indicators for the inferencing platform.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 3–5+ years of experience in Site Reliability Engineering, DevOps, or a related role supporting large-scale customer-facing services in a public cloud environment.
- Strong programming or scripting skills in Python, Go, or Java.
- Experience with Docker and Kubernetes.
- Experience with monitoring and observability tools such as Prometheus, Grafana, ELK Stack, and Datadog.
- Experience with Infrastructure as Code tools such as Terraform and CloudFormation.
- Familiarity with CI/CD tools such as Jenkins, GitHub Actions, and ArgoCD.
- Experience troubleshooting complex distributed systems.
- Experience with hybrid cloud and on-premises or data-center infrastructure is preferred.
- Production experience supporting ML/AI inferencing services is preferred.
- Familiarity with GPU-accelerated computing, NVIDIA GPUs, vLLM, SGLang, Ray, and MLOps is preferred.
- Experience managing and tuning SQL or NoSQL databases and Redis or Memcached is preferred.
- Strong Linux/Unix system administration fundamentals are preferred.
Benefits
- Flexible work environment
- Competitive benefits including medical, dental, vision, disability, life, AD&D, and flexible spending account plans
- Health Savings Account with employer contribution
- Well-being benefits including Headspace, Gympass+, One Medical, and counseling services
- Equity is included in the compensation package
Tech Stack
Categories
About SambaNova Systems
Welcome to SambaNova: Revolutionizing AI Capacity At SambaNova, we're empowering developers, enterprises, governments, and data centers to unlock their full AI potential. Our full-stack infrastructure, from chips to models, enables lightning-fast performance, low power consumption, and high-efficiency computing. Our Mission To give every developer, enterprise, government and data center absolute sovereignty over their own data, models and AI infrastructure – to future-proof the AI workloads that will power and scale tomorrow. Our Technology We give our customers the optionality to experience SambaNova through the cloud or on-premise. Samba Cloud delivers the fastest inferences on the largest open source models like Llama 4 and DeepSeek. Developers can get started building in minutes with our OpenAI compatible APIs. All customers start on the developer tier and when they need more capacity can scale into our enterprise tier. SambaStack is our on-premise offering which includes the system, the platform, and foundation models. These components combine into a powerful technology stack that delivers unparalleled performance, ease of use, accuracy, data privacy, and the ability to power every use case across the world's largest organizations. SambaManaged is a modular and ready-to-deploy AI cloud designed to deliver unmatched efficiency for data centers and cloud service providers. This solution allows organizations to quickly deploy advanced AI inference services—without the need for costly infrastructure upgrades or specialized expertise—in as little as 90 days. At the heart of SambaNova innovation is the Reconfigurable Dataflow Unit (RDU). Purpose built for AI workloads, the RDU takes advantage of a dataflow architecture and a three-tiered memory design. The three tiers of memory enable the platform to run hundreds of models on a single node and to switch between them in microseconds. In 2023, SambaNova released its 4th generation RDU chip, the SN40L.