CloudZero

Senior CloudOps Engineer

CloudZero
Apply
2 months ago
Boston, MA, USA or San Francisco, CA, USASenior

Base Salary

$130k - $190k/yr

Responsibilities

  • Own reliability for CloudZero’s real-time ingestion path, including cross-team SLOs, failure modes, and resulting architectural changes.
  • Define and enforce reliability standards for shared critical paths, including error-budget decisions before launch.
  • Instrument systems and build observability so failures surface quickly and debugging is data-driven.
  • Build production Python reliability tooling, including load generators, fault-injection harnesses, SLO libraries, deployment safety checks, services, automation, and agents.
  • Design and maintain CloudFormation and SAM modules for reliable and cost-efficient cloud resources.
  • Automate deployments, scaling, backups, limit changes, and other repetitive operational work.
  • Evaluate and improve autonomous agents and make services legible to AI tooling through ownership and service metadata.
  • Partner with product engineering on resilient service design, operational architecture reviews, deployment pipelines, shared templates, and SLO adoption.
  • Optimize infrastructure for cost and performance and drive adoption of reliability practices across more than 40 engineers.

Requirements

  • Strong production Python experience as a primary language, including ownership, testing, and maintenance at scale.
  • Experience defining SLOs, including deliberate decisions about what not to alert on.
  • Experience operating asynchronous, event-driven systems and reasoning about back-pressure, consumer lag, replay, poison messages, and partial failure using technologies such as Kafka, Kinesis, SQS, Pulsar, or Step Functions.
  • At least 5 years building and operating distributed systems in AWS, with ownership of reliability outcomes.
  • Practical Infrastructure as Code experience with CloudFormation and SAM, or transferable depth in Terraform or Pulumi.
  • Hands-on experience instrumenting systems with monitoring tools such as Sumo Logic, Datadog, Prometheus, or Splunk.
  • Proven ability to debug production issues under pressure and drive reliability or platform changes across teams without direct authority.
  • Experience with or appetite for AI models and tooling such as Claude, Codex, or Gemini.
  • Strong system design, documentation, communication, and cross-functional collaboration skills.
  • Preferred experience includes chaos engineering, load testing, internal developer portals such as Cortex or Backstage, test automation, ephemeral test environments, GitHub Actions at scale, or LLM-backed tooling.

Tech Stack

Apache KafkaAWSAzureDatadogGitHub ActionsGoogle Cloud PlatformPrometheusPythonSplunkSumo LogicTerraform

Categories

DevOpsSite Reliability
CloudZero

About CloudZero

51-200 employees

CloudZero builds a SaaS FinOps platform that analyzes cloud and AI infrastructure spend to provide unit economics and engineering-level cost insights. It helps product, finance, and DevOps teams allocate costs, detect anomalies, and optimize usage across providers such as AWS. Founded in 2016 and headquartered in Boston, the privately held company sells subscriptions to enterprises building software at scale.

Contact me