7 days ago
Base Salary
$148k - $220k/yr
Responsibilities
- Develop automation for deployment, packaging, monitoring visibility, and efficient cloud operations.
- Collaborate with infrastructure engineers and developers to improve deployment performance, reliability, scalability, and automation.
- Debug and troubleshoot service bottlenecks and provide advanced tier 2 and tier 3 support for cloud data services.
- Monitor and analyze system health, availability, latency, and performance across servers, containers, databases, and backend infrastructure.
- Respond to complex production incidents, perform root cause analysis, and resolve cross-platform issues involving operating systems, networking, and databases.
- Create system documentation and runbooks and maintain accessible critical system information.
- Identify, diagnose, and resolve complex security issues while staying current with security protocols.
- Use Atlassian and cloud service management tools to track and resolve issues by priority.
- Participate in a rotation-based 24x7 on-call schedule, including weekends and non-traditional work hours.
Requirements
- At least 5 years of experience in scripting and infrastructure automation.
- Experience with PowerShell, Python, Go, or Ruby.
- Deep knowledge of containers, Kubernetes, serverless computing implementation, and distributed systems design patterns.
- Proficiency with Linux/Unix and CoreOS.
- Experience with AWS, Azure, or Google Cloud.
- Knowledge of DevOps and Site Reliability Engineering development methodologies.
- Ability to lead a scrum team, influence stakeholders, maintain a product backlog, and manage sprints.
- Bachelor of Science in Computer Science, a master’s degree, or equivalent experience.
Benefits
- Comprehensive benefits package that may include health insurance, life insurance, retirement or pension plans, paid time off, leave options, performance-based incentives, employee stock purchase plans, and restricted stock units.
- Benefits and offerings vary by country and region and are subject to local laws, regulations, and company policies.
- The role includes a global rotation-based 24x7 on-call schedule, with every-other-week rotations and non-traditional workdays and hours including weekends.
Tech Stack
Categories
DevOpsSite Reliability
About NetApp
NetApp builds data storage and data management software and hardware for enterprises, including ONTAP-powered all-flash arrays and hybrid cloud services that integrate with AWS, Azure, and Google Cloud. It sells systems and subscriptions that support backup, disaster recovery, databases, and container/Kubernetes workloads. Founded in 1992 and headquartered in San Jose, California, NetApp is a public company listed on NASDAQ.
