
Senior Staff Software Engineer – SRE & AIOps
ServiceNow5 hours ago
Base Salary
$191k - $334k/yr
Responsibilities
- Design, deploy, and operate enterprise-scale Kubernetes clusters across hybrid and multi-cloud environments.
- Architect closed-loop auto-remediation systems using agentic AI and machine learning to detect, predict, and resolve infrastructure failures.
- Evolve monitoring, incident management, log aggregation, and observability tooling for global follow-the-sun operations.
- Establish SLOs, error budgets, alerting policies, automated runbooks, and playbooks that improve incident response and reduce alert fatigue.
- Design Infrastructure-as-Code frameworks and GitOps pipelines for reproducible, auditable infrastructure deployments.
- Architect hybrid cloud and data center operations, including workload migrations, disaster recovery, edge computing, and cost optimization.
- Drive adoption of containerization, microservices, DevOps, CI/CD, service mesh, and network security patterns.
- Design on-call rotations, escalation policies, incident command systems, and post-incident review processes for global teams.
- Mentor SRE, infrastructure, and DevOps engineers on reliability, incident investigation, automation, and AI applications.
- Reduce operational toil through automation of provisioning, incident response, and cost optimization.
Requirements
- 12+ years of software engineering or infrastructure operations experience, including 7+ years in senior SRE, DevOps, or cloud platform engineering roles with a Bachelor's degree; alternative paths include 8 years with a Master's degree, 5 years with a PhD, or equivalent experience.
- 5+ years of hands-on experience designing, deploying, and operating production Kubernetes clusters at enterprise scale.
- Expertise with Infrastructure-as-Code tools such as Terraform, CloudFormation, or equivalent platforms.
- Demonstrable experience across at least two of AWS, Azure, and GCP, including relevant compute, networking, storage, and observability services.
- Experience designing or operating 24/7 follow-the-sun on-call models, escalation policies, runbooks, and incident response.
- Proven ability to architect automated remediation systems, including AI-driven closed-loop systems that reduce manual toil.
- Strong Linux administration, performance troubleshooting, and scripting experience with Python, Go, and Bash.
- Deep understanding of distributed systems, eventual consistency, cascading failures, network partitions, and Byzantine fault tolerance.
- Experience applying machine learning and AI-driven insights to anomaly detection, predictive alerting, and intelligent remediation.
- Bachelor's degree in computer science, computer engineering, or a related field, or equivalent professional experience.
- Preferred qualifications include Kubernetes certification such as CKA or CKAD, AI/ML certification or agentic AI expertise, service mesh or advanced Kubernetes networking experience, cloud migration experience, hybrid-cloud cost optimization experience, and infrastructure team leadership or mentoring.
Benefits
- Base pay is $190,900-$334,100, with equity when applicable, variable or incentive compensation, and benefits.
- Health plans, flexible spending accounts, a 401(k) plan with company match, ESPP, matching donations, flexible time away, and family leave programs are offered.
- The role uses a flexible work persona, with work arrangements assigned based on role and location.
- ServiceNow offers an accessible and inclusive application process with reasonable accommodations.
Categories
DevOpsSite Reliability
About ServiceNow
ServiceNow (NYSE: NOW) makes the world work better for everyone. Our cloud-based platform and solutions help digitize and unify organizations so that they can find smarter, faster, better ways to make work flow. So employees and customers can be more connected, more innovative, and more agile. And we can all create the future we imagine. The world works with ServiceNow. For more information, visit www.servicenow.com.