Gen Digital

Principal Site Reliability Engineer

Gen Digital
Apply
9 hours ago
Tempe, AZ, USAStaff+
H1B sponsor

Base Salary

$150k - $160k/yr

Responsibilities

  • Define and lead the platform’s long-term reliability strategy, technical direction, and operational standards.
  • Lead architecture decisions for highly available, fault-tolerant distributed systems running on AWS, GCP, Kubernetes, and GKE.
  • Establish SRE practices including SLIs, SLOs, error budgets, production readiness reviews, service ownership standards, and operational risk management.
  • Guide platform architecture, deployment patterns, and infrastructure automation using Terraform and infrastructure-as-code tooling.
  • Design and implement observability capabilities covering metrics, logging, tracing, alerting, and dashboards.
  • Lead major incident response, improve escalation and response processes, and ensure corrective actions through post-incident reviews.
  • Set engineering guardrails, review RFCs, and provide technical leadership on high-impact reliability and architecture initiatives.
  • Drive resilience engineering, disaster recovery, failover design, capacity forecasting, and business continuity planning.
  • Partner with security and compliance stakeholders on infrastructure security and operational controls for PCI and SOC 2 environments.
  • Improve CI/CD systems and software delivery workflows to increase deployment safety, speed, repeatability, and developer experience.
  • Reduce operational toil through automation, self-service platform capabilities, and better engineering abstractions.
  • Mentor senior engineers and technical leads and align global SRE teams on shared standards and practices.
  • Advise engineering and product leadership on reliability tradeoffs, infrastructure investments, and operational risk.

Requirements

  • 8+ years of experience in SRE, DevOps, platform engineering, or infrastructure engineering, including leadership of large-scale cloud initiatives in production.
  • Deep expertise in AWS and Kubernetes and experience designing and operating large-scale containerized microservice-based systems.
  • Experience with infrastructure as code, especially Terraform, and AWS services and container platforms such as EKS, ECS, or GKE.
  • Proven experience implementing SLIs, SLOs, error budgets, incident management, observability, and production readiness standards.
  • Strong understanding of distributed-systems reliability, performance, scaling, availability engineering, and failure-mode analysis.
  • Experience leading complex cross-team technical initiatives and influencing architecture and engineering practices beyond direct reporting lines.
  • Strong CI/CD and delivery-engineering background, including Jenkins, CircleCI, or GitHub Actions.
  • Proficiency in at least one programming language such as Python or Go for automation, tooling, and production-quality code.
  • Experience in security-conscious and compliant environments with familiarity with PCI, SOC 2, and NIST.
  • Excellent written and verbal communication skills for guiding decisions from senior engineers through executive stakeholders.
  • Experience using AI-assisted and agentic engineering tools to improve productivity, automation, operational insight, and developer workflows.
  • Ability to balance strategic direction with hands-on execution in high-scale, high-ownership environments.

Tech Stack

Categories

Site Reliability
Gen Digital

About Gen Digital

1,001-5,000 employees

Gen Digital is a public consumer cyber safety company that sells subscription software and services for cybersecurity, online privacy, and identity protection under brands including Norton, Avast, LifeLock, Avira, AVG, CCleaner, and ReputationDefender. Formed in 2022 by the merger of NortonLifeLock and Avast, it is headquartered in Tempe, Arizona, and serves more than 150 countries with a user base approaching 500 million. Revenue comes primarily from direct-to-consumer subscriptions and partnerships.

Contact me