7 months ago
Atlanta, GA, USASenior
Responsibilities
- Define and drive the reliability strategy for Saviynt’s federal SaaS platform.
- Design, build, and maintain shared infrastructure services and internal platform tools.
- Develop infrastructure provisioning and management automation primarily in Go or Python.
- Architect reusable cloud patterns and modules across AWS, Azure, and other cloud environments.
- Build shared event-driven architecture, messaging, distributed systems, and service mesh components.
- Manage multi-region infrastructure and highly available platform services.
- Establish centralized observability and monitoring platforms for engineering teams.
- Design well-documented RESTful APIs for internal infrastructure services.
- Design and manage highly available relational database services and shared data platforms.
- Collaborate with product engineering teams and provide technical guidance and support.
- Manage vulnerability-related responsibilities and help teams meet customer-facing SLAs.
- Design continuous delivery processes for government deployments and participate in on-call rotations.
Requirements
- At least 1 year of experience as a Senior SRE focused on building tools and services for engineers.
- Deep production experience with Kubernetes, including single-tenant and multi-tenant platform architectures.
- Strong Go and Python programming skills for backend services and infrastructure automation.
- Hands-on experience with at least one major cloud provider: AWS, GCP, or Azure; multi-cloud experience is preferred.
- Experience designing event-driven architectures and shared message queuing services.
- Experience with CI/CD pipeline tools, especially GitLab CI, and automated delivery processes.
- Experience designing and operating distributed systems and reliable shared components.
- Experience with multi-region cloud environments and globally available platforms.
- Proficiency with observability and monitoring platforms such as Prometheus, Grafana, ELK, or Datadog.
- Strong RESTful API design and implementation experience.
- Knowledge of service mesh concepts and practical experience with solutions such as Istio.
- Hands-on experience managing relational databases such as MySQL or PostgreSQL as a service.
- Strong communication skills and a customer-centric approach to supporting internal engineering teams.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical or military experience.
- FedRAMP compliance, government security requirements, and secure CI/CD experience in restricted or regulated environments are preferred.
Benefits
- Work on a mission-critical SaaS platform used by global enterprises and government institutions.
- Solve complex reliability challenges at scale and influence company-wide architecture and engineering culture.
- Competitive compensation, benefits, and growth opportunities.
- Participate in an on-call rotation supporting critical shared infrastructure.
Tech Stack
AmbassadorApache KafkaAWSAzureDatadogGitLab CI/CDGoGoogle Cloud PlatformGrafanaIstioKubernetesMySQLPostgreSQLPrometheusPython
Categories
DevOpsSite Reliability
