5 months ago
Bengaluru, IndiaSenior
Responsibilities
- Implement and manage cloud-native systems on AWS using automation and best-in-class tools.
- Operate and enhance Kubernetes clusters, deployment pipelines, and service meshes.
- Design, build, and maintain reliable, secure, and scalable infrastructure for a multi-tenant SaaS platform.
- Define and maintain SLOs, SLAs, and error budgets while addressing availability and performance issues.
- Develop infrastructure-as-code for repeatable and auditable provisioning.
- Build internal platform tools and automation for provisioning, monitoring, and operational efficiency.
- Monitor infrastructure and applications and improve user experience reliability.
- Participate in shared on-call rotations, troubleshoot outages, and coordinate incident responses as Incident Commander.
- Improve incident response processes and reduce Mean Time to Detect and Mean Time to Recovery.
- Diagnose complex production issues independently and work with software engineers on observability, fault tolerance, and reliability.
- Automate runbooks, health checks, alerting, testing, canary deployments, and rollback strategies.
- Contribute to security best practices, compliance automation, chaos engineering, resilience testing, and cost optimization.
Requirements
- Bachelor’s degree in Computer Science or equivalent practical experience.
- At least 5 years of experience as a Site Reliability Engineer or Platform Engineer.
- Strong knowledge of software development best practices and programming or scripting with Python, Go, or Bash.
- Hands-on experience with AWS, GCP, or Azure and supporting SaaS products.
- Experience with Terraform, Kubernetes, Docker, Helm, and container-based deployment ecosystems.
- Experience supporting multi-tenant microservices architectures and CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, FluxCD, or Crossplane.
- Experience managing monitoring solutions such as Datadog.
- Experience with production operations, rotating on-call schedules, critical incident management, and post-incident reviews.
- Deep understanding of Linux systems, networking, TCP/IP, VPNs, cloud architectures, service meshes, and storage systems.
- Experience with VPCs, subnets, firewalls, load balancers, Istio, S3, and EBS.
- Understanding of zero-downtime, blue/green, and canary deployment strategies.
- Exposure to SOC 2, ISO 27001, or HIPAA; FedRAMP experience is a plus.
- Strong troubleshooting, problem-solving, communication, collaboration, organization, and English-language skills.
- Experience with the listed tools is described as beneficial but not mandatory, and candidates are encouraged to apply without meeting every requirement.
Benefits
- Hybrid work arrangement, as indicated by the #LI-Hybrid designation.
- Opportunity to work with a global team across more than 75 nationalities and 9 offices.
- Work on digital employee experience software used by more than 1,300 customers and 18 million employees.
Tech Stack
AWSAzureBashDatadogDockerGitHub ActionsGitLab CI/CDGoGoogle Cloud PlatformHelmIstioJenkinsKubernetesLinuxPythonTerraform
Categories
DevOpsSite Reliability
