ServiceNow

Staff Software Engineer - SRE & AIOps

ServiceNow
Apply
13 hours ago
Remote, United States or Atlanta, GA, USAStaff+
H1B sponsor

Responsibilities

  • Deploy, operate, and troubleshoot production Kubernetes clusters across hybrid and multi-cloud environments.
  • Design and maintain automated remediation, self-healing, alerting, and runbook systems to reduce MTTR and operational toil.
  • Develop SRE tooling involving monitoring, incident management, log aggregation, and observability integrations.
  • Build Infrastructure-as-Code frameworks and GitOps pipelines for reproducible infrastructure deployments with security and compliance guardrails.
  • Support cloud, on-premises, data center, and multi-region operations, including workload and cost optimization.
  • Participate in on-call rotations, incident response, escalation processes, and post-incident reviews.
  • Promote containerization, microservices, DevOps, CI/CD, network security, and reliability engineering practices.
  • Mentor junior SRE engineers and contribute to documentation, knowledge sharing, and continuous improvement.

Requirements

  • 8+ years of software engineering or infrastructure operations experience with a bachelor's degree, 6+ years with a master's degree, 3+ years with a PhD, or equivalent work experience.
  • At least 2 years of hands-on experience operating production Kubernetes clusters.
  • Hands-on experience with at least one major cloud platform: AWS, Azure, or GCP.
  • Proficiency with at least one Infrastructure-as-Code tool, including Terraform or CloudFormation.
  • Experience with automated remediation, alert systems, on-call operations, incident response, runbooks, and 24/7 operational models.
  • Strong Linux system administration, performance troubleshooting, and scripting skills using Python, Go, or Bash.
  • Understanding of distributed systems, fault tolerance, resilience patterns, hybrid cloud operations, and reliability engineering.
  • Bachelor's degree in computer science, computer engineering, or a related field, or equivalent professional experience.
  • Preferred qualifications include Kubernetes certification such as CKA or CKAD, service mesh or advanced Kubernetes networking experience, cloud migration or modernization experience, and cloud cost optimization experience.
  • Experience integrating AI into work processes, decision-making, or problem-solving is required.

Benefits

  • Flexible or remote work persona, depending on assigned location and role eligibility.
  • Regular employee position in the North America and Canada region.
  • Opportunity to work on infrastructure automation and reliability engineering at global scale with a path toward senior technical leadership.
ServiceNow

About ServiceNow

10,000+ employees

ServiceNow builds a cloud platform for enterprise digital workflows, covering IT service management, customer service, HR service delivery, security operations, and operations management, plus tools for custom app development. It sells subscription SaaS to large organizations and public-sector agencies to automate processes and connect data across systems. Founded in 2004 and headquartered in Santa Clara, California, ServiceNow is a public company listed on the NYSE under the ticker NOW.

Contact me