7 days ago
Base Salary
$228k - $339k/yr
Responsibilities
- Build automation that improves operational efficiency and reduces risk.
- Improve reliability, scalability, and deployment velocity for cloud services.
- Troubleshoot complex production issues across infrastructure and application layers.
- Lead incident response and root cause analysis for critical service events.
- Monitor and optimize service health, availability, latency, and performance.
- Operate Kubernetes- and container-based cloud infrastructure.
- Provide advanced tier 2/3 support for cloud data services.
- Strengthen security posture and resolve complex security issues.
- Create and maintain runbooks and operational documentation.
- Influence architecture and implementation decisions using production insights.
Requirements
- 12+ years of experience in software, infrastructure, or site reliability engineering.
- Full-stack engineering experience is required.
- Strong coding and scripting skills in Python, Go, or PowerShell.
- Deep experience with Linux, Kubernetes, containers, and distributed systems.
- Experience with Microsoft Azure, AWS, or Google Cloud.
- Experience using AI to accelerate automation development and optimize processes.
- Ability to derive actionable business intelligence from structured and unstructured logs, events, traces, and metrics using AI and MCP-enabled agentic workflows.
- Strong understanding of DevOps and site reliability engineering practices.
- Excellent troubleshooting skills across infrastructure and application layers.
Benefits
- Comprehensive benefits package including health insurance, life insurance, retirement or pension plans, paid time off, leave options, performance-based incentives, employee stock purchase plans, and/or restricted stock units, subject to regional variations and local policies.
About NetApp
Build an intelligent data infrastructure with NetApp that brings it all together — a smarter way to let data thrive. Any application, any data, anywhere.