
Staff Software Engineer – SRE & AIOps
ServiceNow5 hours ago
Responsibilities
- Design, deploy, and operate production Kubernetes clusters across hybrid and multi-cloud environments.
- Architect closed-loop auto-remediation systems using anomaly detection, machine learning, agentic AI, and automated runbooks.
- Build and evolve monitoring, incident management, log aggregation, and observability tooling.
- Establish SLOs, error budgets, alerting policies, escalation procedures, and automated incident-response playbooks.
- Develop Infrastructure-as-Code frameworks and GitOps pipelines for reproducible and auditable infrastructure deployments.
- Architect hybrid-cloud, data-center, edge, workload migration, disaster-recovery, and cost-optimization solutions.
- Drive containerization, microservices, service mesh, network security, and DevOps practices across engineering teams.
- Design global on-call rotations and incident command processes, and lead post-incident reviews.
- Mentor SRE and infrastructure engineers on reliability, automation, incident investigation, and AI applications.
- Reduce operational toil through automation of infrastructure provisioning, incident response, and cost optimization.
Requirements
- 8+ years of software engineering or infrastructure operations experience, including 5+ years in SRE, DevOps, or cloud platform engineering, with the stated degree-based alternatives or equivalent experience.
- 4+ years of hands-on experience designing, deploying, and operating production Kubernetes clusters at scale.
- Strong experience with Kubernetes cluster design, orchestration, resource management, networking, security controls, and troubleshooting.
- Hands-on experience across at least two of AWS, Azure, and GCP, including multi-region cloud architecture.
- Experience with Infrastructure-as-Code tools such as Terraform or CloudFormation and GitOps platforms.
- Strong Linux administration, systems troubleshooting, and scripting experience with Python, Go, or Bash.
- Experience designing automated remediation, anomaly detection, alert correlation, runbook automation, and self-healing systems.
- Experience with 24/7 follow-the-sun on-call operations, incident response, escalation policies, and runbook development.
- Experience managing on-premises and public-cloud infrastructure, hybrid networking, disaster recovery, and workload migration.
- Understanding of distributed systems, eventual consistency, cascading failures, and network partitions.
- Ability to apply AI and machine-learning insights to infrastructure operations.
- Bachelor's degree in computer science, Computer Engineering, or a related field, or equivalent professional experience.
- Preferred qualifications include CKA or CKAD certification, service-mesh or advanced Kubernetes networking experience, cloud migration experience, hybrid-cloud cost optimization, and infrastructure-team mentoring.
Benefits
- Base pay of C$125,700–C$220,000, plus equity when applicable, variable/incentive compensation, and benefits.
- Health plans, flexible spending accounts, a 401(k) plan with company match, ESPP, matching donations, flexible time away, and family leave programs.
- Regular full-time employee position with a remote work persona in North America and Canada.
- Flexible distributed-work arrangements based on assigned work location and work persona.
- Reasonable accommodations and an inclusive equal-opportunity workplace.
Categories
DevOpsSite Reliability
About ServiceNow
ServiceNow (NYSE: NOW) makes the world work better for everyone. Our cloud-based platform and solutions help digitize and unify organizations so that they can find smarter, faster, better ways to make work flow. So employees and customers can be more connected, more innovative, and more agile. And we can all create the future we imagine. The world works with ServiceNow. For more information, visit www.servicenow.com.