Nexthink

Senior Site Reliability Engineer

Nexthink
Apply
5 months ago
Bengaluru, IndiaSenior

Responsibilities

  • Implement and manage cloud-native systems on AWS using automation and best-in-class tools.
  • Operate and enhance Kubernetes clusters, deployment pipelines, and service meshes.
  • Design, build, and maintain reliable, secure, and scalable infrastructure for a multi-tenant SaaS platform.
  • Define and maintain SLOs, SLAs, and error budgets while addressing availability and performance issues.
  • Develop infrastructure-as-code for repeatable and auditable provisioning.
  • Build internal platform tools and automation for provisioning, monitoring, and operational efficiency.
  • Monitor infrastructure and applications and improve user experience reliability.
  • Participate in shared on-call rotations, troubleshoot outages, and coordinate incident responses as Incident Commander.
  • Improve incident response processes and reduce Mean Time to Detect and Mean Time to Recovery.
  • Diagnose complex production issues independently and work with software engineers on observability, fault tolerance, and reliability.
  • Automate runbooks, health checks, alerting, testing, canary deployments, and rollback strategies.
  • Contribute to security best practices, compliance automation, chaos engineering, resilience testing, and cost optimization.

Requirements

  • Bachelor’s degree in Computer Science or equivalent practical experience.
  • At least 5 years of experience as a Site Reliability Engineer or Platform Engineer.
  • Strong knowledge of software development best practices and programming or scripting with Python, Go, or Bash.
  • Hands-on experience with AWS, GCP, or Azure and supporting SaaS products.
  • Experience with Terraform, Kubernetes, Docker, Helm, and container-based deployment ecosystems.
  • Experience supporting multi-tenant microservices architectures and CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, FluxCD, or Crossplane.
  • Experience managing monitoring solutions such as Datadog.
  • Experience with production operations, rotating on-call schedules, critical incident management, and post-incident reviews.
  • Deep understanding of Linux systems, networking, TCP/IP, VPNs, cloud architectures, service meshes, and storage systems.
  • Experience with VPCs, subnets, firewalls, load balancers, Istio, S3, and EBS.
  • Understanding of zero-downtime, blue/green, and canary deployment strategies.
  • Exposure to SOC 2, ISO 27001, or HIPAA; FedRAMP experience is a plus.
  • Strong troubleshooting, problem-solving, communication, collaboration, organization, and English-language skills.
  • Experience with the listed tools is described as beneficial but not mandatory, and candidates are encouraged to apply without meeting every requirement.

Benefits

  • Hybrid work arrangement, as indicated by the #LI-Hybrid designation.
  • Opportunity to work with a global team across more than 75 nationalities and 9 offices.
  • Work on digital employee experience software used by more than 1,300 customers and 18 million employees.

Tech Stack

AWSAzureBashDatadogDockerGitHub ActionsGitLab CI/CDGoGoogle Cloud PlatformHelmIstioJenkinsKubernetesLinuxPythonTerraform

Categories

DevOpsSite Reliability
Nexthink

About Nexthink

1,001-5,000 employees
Contact me