1 day ago
Base Salary
$80k - $148k/yr
Responsibilities
- Monitor system health, performance metrics, and alerts; troubleshoot issues and restore services during incidents.
- Lead incident response, root-cause analysis, post-incident reviews, and preventive remediation.
- Design and maintain automation tools, scripts, self-healing capabilities, deployment processes, and infrastructure-as-code practices.
- Optimize system resources, performance, scalability, availability, and operational efficiency.
- Plan capacity growth and ensure systems are resilient, highly available, and scalable.
- Maintain configuration consistency and enforce standards across environments.
- Collaborate with development, operations, security, and other stakeholders on reliability goals, security controls, and compliance.
- Perform other related tasks as required.
Requirements
- Seasoned hands-on experience in Site Reliability Engineering or related roles, designing and maintaining highly available and scalable systems.
- Seasoned expertise in Linux/Unix, networking, system administration, cloud platforms, infrastructure architectures, and fault-tolerant design.
- Proficiency in programming and scripting languages including Python, Go, Java, Ruby, Bash, or PowerShell.
- Experience with infrastructure-as-code tools such as Terraform or CloudFormation and containerization technologies such as Docker or Kubernetes.
- Experience designing automation frameworks, CI/CD pipelines, and deployment strategies using tools such as Jenkins, GitLab CI/CD, or CircleCI.
- Experience with performance monitoring, optimization, troubleshooting, incident management, root-cause analysis, and post-incident reviews.
- Understanding of security principles, security controls, vulnerability assessment, compliance requirements, DevOps principles, and Agile methodologies.
- Strong problem-solving, analytical, communication, collaboration, leadership, and continuous-improvement skills.
- Bachelor's degree or equivalent in Computer Science, Information Technology, or a related field.
- Relevant certifications such as AWS Certified DevOps Engineer - Professional, Google Cloud Professional DevOps Engineer, or Certified Kubernetes Administrator preferred.
Benefits
- Remote working arrangement.
- U.S. annual starting pay range of $80,000.00 - 114,000.00 - 148,000.00, with actual compensation dependent on location, experience, technical skills, and other qualifications.
- Equal opportunity workplace committed to diversity and inclusion.
