8 hours ago
Base Salary
$153k - $205k/yr
Responsibilities
- Design, build, secure, troubleshoot, and operate highly available Kubernetes platforms across hybrid and public-cloud environments.
- Create reusable Terraform modules and automated, reviewable infrastructure delivery workflows.
- Develop maintainable backend services, internal tools, and operational automation in Go, Python, or JavaScript/TypeScript.
- Partner with platform, product, application engineering, and Security teams to design reliable, performant, secure, and cost-effective solutions.
- Improve CI/CD, deployment automation, progressive delivery, and operational ownership across the production lifecycle.
- Define observability practices covering metrics, logs, traces, alerting, and dashboards.
- Participate in on-call, lead incident response, perform root-cause analysis, and drive blameless postmortems and corrective actions.
- Establish and operate SLIs, SLOs, error budgets, capacity plans, disaster-recovery tests, and resilience improvements.
- Apply AI-assisted and data-driven operational techniques to improve detection, reduce alert noise, and accelerate troubleshooting.
- Contribute through code reviews, documentation, knowledge sharing, mentoring, and team development.
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a closely related software engineering role supporting production systems.
- Deep hands-on experience designing, operating, securing, and troubleshooting Kubernetes clusters and containerized workloads at scale.
- Strong Terraform experience, including reusable modules, state and environment management, and automated infrastructure delivery.
- Production software-development experience in Go, Python, or JavaScript/TypeScript for backend services, tooling, and automation.
- Demonstrated success improving the reliability, performance, scalability, or cost efficiency of distributed production systems.
- Experience with cloud infrastructure, networking, IAM, DNS, load balancing, routing, service networking, and secure connectivity.
- Strong observability and troubleshooting skills involving metrics, logs, traces, alerting, and incident data.
- Experience with SLIs, SLOs, error budgets, incident management, postmortems, and disaster-recovery practices.
- Familiarity with CI/CD, GitOps or deployment automation, and canary or blue-green deployment strategies.
- Security-minded approach and experience partnering with Security and engineering teams in regulated or high-availability environments.
- Clear written and verbal communication, strong ownership, and sound judgment balancing speed, risk, and operational excellence.
- Experience applying AI-assisted tooling to engineering or operations workflows is a plus.
Benefits
- Remote work arrangement, indicated by the #LI-Remote designation.
- Base compensation range of $152,500-$205,000, with starting pay determined by experience, skills, qualifications, and organizational factors.
- Inclusive equal-opportunity workplace with interview accommodations available for candidates with disabilities.
Categories
DevOpsSite Reliability
