6 months ago
Remote, United StatesSenior
Responsibilities
- Own the reliability and performance of cloud infrastructure across AWS and GCP.
- Manage and optimize Kubernetes clusters and container orchestration.
- Drive database reliability engineering, including performance tuning and scaling.
- Build and maintain CI/CD pipelines for rapid, safe deployments.
- Run incident response and on-call rotations.
- Partner with product engineers to design scalable and resilient systems.
Requirements
- Strong AWS experience, particularly with ECS, Aurora, and CloudWatch.
- Experience with GCP as the company expands across clouds.
- Expertise with Kubernetes and container orchestration.
- Database reliability engineering experience, including performance tuning and scaling.
- Experience owning CI/CD pipelines and handling incident response.
- Background at a B2B SaaS company serving enterprise customers, ideally in infrastructure.
- Experience deploying and supporting on-premises or hybrid environments is a bonus.
- Python backend familiarity is a bonus.
- Experience at an early-stage or high-growth company is a bonus.
Benefits
- Four weeks of paid vacation, paid sick leave, and paid parental leave.
- Professional development budget for conferences, courses, and certifications.
- Choice of laptop and accessories.
- Comprehensive health, dental, and vision coverage.
- Opportunities to interact directly with enterprise customers.
- The role is at a 25-person, mostly engineering, early-stage company.
Tech Stack
Categories
DevOpsSite Reliability
About Runlayer
Runlayer builds an enterprise MCP platform that securely connects AI agents, skills, and tools to internal systems with governance, permissions, and observability. It sells its platform to enterprises adopting AI across functions, focusing on security and control in complex environments. The privately held company is headquartered in New York, with customers including Gusto, Instacart, Opendoor, dbt Labs, and Decagon.
