2 months ago
Atlanta, GA, USA or Milpitas, CA, USAStaff+
Base Salary
$240k - $250k/yr
Responsibilities
- Define and drive the reliability strategy for Saviynt’s SaaS platform across infrastructure, platform, and application teams.
- Design and build reusable shared infrastructure services and platform components for internal development teams.
- Architect, implement, and manage highly available single-tenant and multi-tenant Kubernetes platforms as a service.
- Develop Go-based internal tools and automation for infrastructure provisioning and management.
- Build multi-cloud, multi-region infrastructure patterns and modules across GCP, AWS, and Azure.
- Design shared event-driven architecture, messaging, distributed systems, CI/CD, observability, service mesh, API, and relational database platforms.
- Provide technical guidance and support to product development teams using shared infrastructure services.
- Participate in on-call rotations for critical shared infrastructure.
Requirements
- At least 1 year of experience as a Principal SRE focused on building tools and services for other engineers.
- Deep production Kubernetes expertise, especially with platform-as-a-service and single-tenant or multi-tenant architectures.
- Strong Go and Python programming skills for backend services and infrastructure automation.
- Extensive experience with at least one major cloud provider, with GCP required; multi-cloud abstraction experience is preferred.
- Experience designing event-driven architectures, message queuing systems, distributed systems, CI/CD delivery processes, and globally available multi-region platforms.
- Experience with observability and monitoring platforms, RESTful API design, service mesh solutions, and relational databases.
- Strong communication skills and a customer-centric approach toward internal development teams.
- Advanced Professional GCP Certification is required.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical or military experience is required.
Benefits
- Growth and learning opportunities in a high-growth Platform as a Service company.
- Welcoming and positive work environment with challenging, rewarding work that impacts customers.
- Requires participation in on-call rotations and adherence to information security, privacy, and annual security training policies.
Tech Stack
AmbassadorApache KafkaAWSAzureDatadogGitLab CI/CDGoGoogle CloudGoogle Cloud PlatformGrafanaIstioKubernetesMySQLPostgreSQLPrometheusPythonRabbitMQ
Categories
DevOpsSite Reliability
