6 hours ago
Responsibilities
- Deploy, operate, and troubleshoot Kubernetes-based services across public cloud, private cloud, and on-premises environments.
- Build and maintain automation for deployments, releases, configuration management, infrastructure, and operational workflows.
- Develop tools and software that improve service reliability, scalability, efficiency, and engineering productivity.
- Collaborate with development, SRE, Platform Engineering, Operations, QA, cloud providers, and other teams to resolve production and deployment issues.
- Improve monitoring, metrics, logging, dashboards, and alerting for application and infrastructure health.
- Investigate complex issues across applications, Kubernetes, networking, infrastructure, and cloud platforms through root cause and resolution.
- Replace repetitive or costly operational processes with reliable automation and tooling.
- Improve scalability, performance, capacity management, operational readiness, runbooks, and engineering practices.
- Participate in production support and the team’s on-call rotation.
- Use AI-assisted development tools and coding agents for automation, tooling, testing, troubleshooting, and operational solutions.
Requirements
- At least 5 years of experience in DevOps, Site Reliability Engineering, Platform Engineering, infrastructure engineering, or a related role; staff-level candidates typically have 8 or more years.
- Strong hands-on Linux experience and production-system troubleshooting skills.
- Experience deploying and operating Kubernetes-based services in production.
- Experience with at least one major public or private cloud platform, such as AWS, GCP, Azure, or OpenStack.
- Understanding of networking fundamentals including TCP/IP, DNS, load balancing, SSL/TLS, and cloud or Kubernetes networking.
- Programming and automation experience using Python, Go, or a similar language.
- Experience with CI/CD, infrastructure automation, or configuration management using Jenkins, Ansible, Terraform, or equivalent technologies.
- Experience with production monitoring and observability using Prometheus, Grafana, Sumo Logic, Elastic Stack, or equivalent platforms.
- Ability to independently own and deliver technical projects and collaborate with geographically distributed teams.
- Hands-on experience using AI-assisted development tools and coding agents, including reviewing, validating, and taking ownership of AI-generated code.
- Preferred experience building internal platforms, developer tooling, release automation, or systems that improve engineering productivity.
- Preferred experience troubleshooting distributed systems and large-scale cloud services, applying strong software engineering practices, and working across multiple cloud or hybrid-cloud environments.
Benefits
- The team supports work across public cloud, private cloud, and on-premises environments.
- Employees receive opportunities for continuous education, mentorship, career growth, and use of modern AI-assisted development tools.
- The role includes participation in a production support and on-call rotation.
Tech Stack
AnsibleAWSAzureGoGoogle Cloud PlatformGrafanaJenkinsKubernetesLinuxOpenStackPrometheusPythonSumo LogicTerraform
Categories
About Netskope
Netskope (NASDAQ: NTSK), a leader in modern security and networking for the cloud and AI era, addresses the needs of both security and networking teams by providing optimized access and real-time, context-based security for the AI ecosystem inclusive of agents, applications, tools, LLMs, people, devices, and data. Thousands of customers, including more than 30 of the Fortune 100, trust the Netskope One platform, its Zero Trust Engine, and its powerful NewEdge network to reduce risk and gain full visibility and control over cloud, AI, SaaS, web, and private applications – providing security and accelerating performance without trade-offs.