Retool

Site Reliability Engineer (SRE)

Retool
Apply
2 months ago

Responsibilities

  • Own reliability across Retool Cloud, managed single-tenant, BYOC, and self-hosted deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations.
  • Build automation for Terraform runs, customer environment updates, upgrade workflows, secret rotations, and migration steps.
  • Improve observability by turning health signals into clear status, likely causes, and recommended actions.
  • Design safer deployment, upgrade, and rollback paths for cloud and managed customers.
  • Create repeatable migration flows toward supported deployment paths such as Blueprints, Kubernetes, and Helm.
  • Partner with product engineers on infrastructure requirements for new Retool products and dependencies.
  • Write documentation, runbooks, design notes, and migration guides for engineers and customers.
  • Lead through ambiguity, make careful risk decisions, and communicate clearly with teams and customers.

Requirements

  • Deep experience operating production infrastructure in AWS.
  • Experience improving reliability for customer-facing SaaS systems.
  • Strong Kubernetes fundamentals.
  • Practical Terraform or infrastructure-as-code experience.
  • Operational judgment around databases, especially Postgres.
  • Experience building or operating observability systems.
  • Programming ability in Go, Python, TypeScript, Java, or Ruby.
  • A strong bias toward automation and eliminating repeated operational work.
  • Clear written communication and comfort working with customer-facing teams or customers.
  • Ability to work effectively in a fast-moving environment with changing priorities and imperfect systems.

Categories

DevOpsSite Reliability
Retool

About Retool

201-500 employees
Contact me