2 months ago
Responsibilities
- Own reliability across Retool Cloud, managed single-tenant, BYOC, and self-hosted deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations.
- Build automation for Terraform runs, customer environment updates, upgrade workflows, secret rotations, and migration steps.
- Improve observability by turning health signals into clear status, likely causes, and recommended actions.
- Design safer deployment, upgrade, and rollback paths for cloud and managed customers.
- Create repeatable migration flows toward supported deployment paths such as Blueprints, Kubernetes, and Helm.
- Partner with product engineers on infrastructure requirements for new Retool products and dependencies.
- Write documentation, runbooks, design notes, and migration guides for engineers and customers.
- Lead through ambiguity, make careful risk decisions, and communicate clearly with teams and customers.
Requirements
- Deep experience operating production infrastructure in AWS.
- Experience improving reliability for customer-facing SaaS systems.
- Strong Kubernetes fundamentals.
- Practical Terraform or infrastructure-as-code experience.
- Operational judgment around databases, especially Postgres.
- Experience building or operating observability systems.
- Programming ability in Go, Python, TypeScript, Java, or Ruby.
- A strong bias toward automation and eliminating repeated operational work.
- Clear written communication and comfort working with customer-facing teams or customers.
- Ability to work effectively in a fast-moving environment with changing priorities and imperfect systems.
Tech Stack
Categories
DevOpsSite Reliability