7 months ago
Atlanta, GA, USA +2 moreSenior / Staff+
Responsibilities
- Design and build reusable shared infrastructure services and platform components for internal development teams.
- Architect, implement, and manage highly available Kubernetes platforms, including single-tenant and multi-tenant deployments.
- Develop infrastructure provisioning, management, and automation tools primarily in Go and Python.
- Build reusable cloud infrastructure patterns across AWS, Azure, and other cloud environments.
- Design shared event-driven architecture, messaging, CI/CD, service mesh, database, API, and observability platforms.
- Operate distributed, multi-region infrastructure with a focus on reliability, performance, fault tolerance, and global availability.
- Provide technical guidance and support to product development teams and treat internal engineers as platform customers.
- Participate in on-call rotations for critical shared infrastructure.
Requirements
- At least 1 year of experience as a Staff SRE focused on building tools and services for other engineers.
- Deep production expertise with Kubernetes, particularly Kubernetes-as-a-platform architectures.
- Strong programming skills in Go and Python, including backend services and infrastructure automation.
- Extensive hands-on experience with AWS, GCP, or Azure; multi-cloud abstraction experience is preferred.
- Experience designing event-driven architectures and shared message queuing systems.
- Experience with CI/CD pipeline tools and establishing automated delivery processes for engineering teams.
- Experience designing and operating distributed systems and reliable shared components.
- Experience with multi-region cloud environments and globally distributed platforms.
- Proficiency with observability and monitoring platforms for shared infrastructure.
- Strong experience designing documented RESTful APIs.
- Knowledge of service mesh concepts and practical experience with platform-oriented service mesh solutions.
- Hands-on experience managing relational databases as a service.
- Excellent communication skills and ability to explain complex technical concepts to technical and non-technical audiences.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical or military experience.
Benefits
- Work on a large-scale, cloud-native SaaS platform.
- Solve complex reliability challenges at scale and influence platform architecture and engineering practices.
- Competitive compensation, benefits, and career growth opportunities.
Tech Stack
AmbassadorApache KafkaAWSAzureDatadogGitLab CI/CDGoGoogle Cloud PlatformGrafanaIstioKubernetesMySQLPostgreSQLPrometheusPython
Categories
DevOpsSite Reliability
