Guidewire Software

Senior Site Reliability Engineer

Guidewire Software
Apply
1 day ago
Bengaluru, IndiaSenior / Staff+

Responsibilities

  • Design, build, and operate highly reliable, scalable infrastructure for a multi-tenant SaaS platform.
  • Automate deployment, provisioning, and operational workflows across cloud infrastructure and applications.
  • Develop internal tools, services, and frameworks that improve engineering efficiency and reduce manual effort.
  • Participate in a 24x7 follow-the-sun on-call rotation for critical production systems.
  • Build platform features, resolve issues, and improve availability, performance, scalability, and resilience.
  • Identify risks, bottlenecks, and failure modes and implement preventative solutions.
  • Build and maintain observability systems, including metrics, logging, tracing, and dashboards.
  • Define and track SLOs and reliability metrics and contribute to incident response, root cause analysis, and blameless postmortems.
  • Design and support secure access patterns using SSO, SAML, and OAuth-based authentication.
  • Collaborate across engineering teams, provide technical guidance, maintain documentation and runbooks, and mentor engineers.

Requirements

  • 8–12 years of hands-on experience in SRE, DevOps, cloud infrastructure, or a related platform engineering role.
  • Strong programming skills in Python or Go; Java and Spring Boot are a plus.
  • Deep experience with AWS and operating production systems at scale.
  • Hands-on expertise with Kubernetes, EKS, Docker, Helm, CNI, and Ingress networking.
  • Experience with Infrastructure as Code tools such as Terraform or Terragrunt.
  • Strong understanding of Linux systems and networking fundamentals.
  • Experience with observability platforms such as Datadog, Prometheus, OpenTelemetry, or CloudWatch.
  • Experience with incident management and production support in a microservices environment.
  • Experience with messaging or streaming systems such as Kafka or SQS and relational databases such as Aurora or RDS is a plus.
  • Working knowledge of SSO, SAML, OAuth, identity providers such as Okta, AWS IAM, VPC security groups, and Kubernetes security primitives.
  • Experience with CI/CD and GitOps tools such as GitHub Actions, TeamCity, Jenkins, FluxCD, or Bitbucket.
  • Bachelor’s degree in Computer Science or a related field, or equivalent experience.
  • AWS or Kubernetes certifications, large-scale SaaS experience, KubeVela or Crossplane exposure, and open-source contributions are preferred.

Benefits

  • Opportunity to work on a mission-critical global SaaS platform used by insurers in more than 40 countries.
  • Opportunity to solve complex infrastructure problems at scale and shape the future of a cloud platform.
  • Collaborative, high-impact engineering culture with opportunities for technical leadership and innovation.
  • The role includes a 24x7 follow-the-sun on-call rotation.
  • Guidewire supports an inclusive workplace and is an equal opportunity and affirmative action employer.

Tech Stack

Apache KafkaAWSDatadogDockerGitHub ActionsGoHelmJavaJenkinsKubernetesLinuxPrometheusPythonSpring BootTeamCityTerraform

Categories

DevOpsSite Reliability
Guidewire Software

About Guidewire Software

1,001-5,000 employees

Guidewire is the platform P&C insurers trust to engage, innovate, and grow efficiently. More than 570 insurers in 43 countries, from new ventures to the largest and most complex in the world, rely on Guidewire products. With core systems leveraging data and analytics, digital, and artificial intelligence, Guidewire defines cloud platform excellence for P&C insurers. We are proud of our unparalleled implementation record, with 1,700+ successful projects supported by the industry’s largest R&D team and SI partner ecosystem. Our marketplace represents the largest partner community in P&C, where customers can access hundreds of applications to accelerate integration, localization, and innovation. For more information, please visit https://www.guidewire.com/.

Contact me