24 hours ago
Remote, United StatesMid Level / Senior
Base Salary
$87k - $123k/yr
Responsibilities
- Own operational excellence for assigned systems and services and support cross-team projects.
- Participate in on-call rotations, respond to incidents, troubleshoot system and deployment issues, and drive resolution.
- Lead postmortems, perform root cause analysis, and implement preventive measures.
- Establish service level indicators, monitoring, alerting, and observability for Kubernetes environments, including EKS.
- Build, maintain, and optimize infrastructure as code across multiple AWS environments.
- Manage and optimize production EKS clusters for availability, resilience, and scalability.
- Support releases and implement scalable services using GitOps and progressive delivery practices.
- Maintain and improve CI/CD pipelines and automation tools to reduce operational toil.
- Lead capacity planning and right-sizing efforts.
- Document systems, runbooks, architecture decisions, and system behaviors, and mentor entry-level SREs.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
- 2–4 years of experience in Site Reliability Engineering, DevOps, or Systems Engineering.
- Experience maintaining high availability and resiliency across AWS services including EKS, EC2, RDS, S3, and VPC.
- Production experience with Kubernetes and Docker, including deployment, troubleshooting, and optimization.
- Proficiency with Terraform, CloudFormation, or similar infrastructure-as-code frameworks.
- Experience with observability platforms and systems and networks supporting incident detection and response.
- Experience with GitLab CI, Jenkins, or equivalent CI/CD tools.
- Knowledge of networking fundamentals, troubleshooting, and high-availability architecture patterns.
- Familiarity with GitOps workflows, incident management, and on-call practices.
- Experience in financial services or regulated industries is preferred.
- Familiarity with SOC 2 or PCI DSS is preferred.
- Experience with Datadog, AppDynamics, New Relic, or similar observability and APM platforms is preferred.
- Strong programming skills in shell, Go, Python, or similar languages are preferred.
- Experience supporting Java Spring Boot applications is preferred.
- Experience with Istio or Linkerd, AWS or Kubernetes certifications, disaster recovery, business continuity planning, and SLO/SLI methodologies is preferred.
Benefits
- Medical, dental, vision, and life insurance.
- 401(k) retirement savings plan with company matching contributions up to 6%, financial advisory services, and investment options.
- Tuition reimbursement up to $5,250 per year.
- Business-casual environment with an option to wear jeans.
- Paid time off upon hire, including paid company holidays and floating holidays.
- 16 hours of paid volunteer time per calendar year.
- Paid parental leave, paid short- and long-term disability, and Family and Medical Leave programs.
- Business Resource Groups supporting inclusion and collaboration.
- Flexible work environment; remote or hybrid employees must provide reliable wired high-speed internet and a suitable home workspace, and may be required to work in the office if requirements are not met.
Categories
Site Reliability
