1 hour ago
Pune, IndiaSenior

Responsibilities

  • Monitor, troubleshoot, and resolve high-priority incidents in live environments.
  • Analyze application logs, diagnose performance bottlenecks, and improve the reliability of core business services.
  • Query, optimize, and debug relational databases using complex SQL scripts.
  • Provision, scale, deploy, monitor, and manage infrastructure across enterprise cloud platforms using infrastructure-as-code.
  • Perform post-mortem failure analysis and build self-healing tools to eliminate repetitive operational work.
  • Participate in on-call incident response and root-cause analysis for complex multi-tier applications.

Requirements

  • Proven experience managing high-availability production environments and handling on-call incident response.
  • Expert-level ability to triage complex multi-tier applications and perform root-cause analysis.
  • Strong capability in writing advanced SQL queries to troubleshoot data layers and performance issues.
  • Hands-on proficiency deploying, monitoring, and scaling core services in enterprise cloud environments.
  • Ability and willingness to write scripts and tools that reduce manual work and operational toil.

Benefits

  • Competitive salary and benefits package.
  • Opportunities for skill development and career growth within a global business.
  • Access to learning, development, and on-the-job experiences.
  • Supportive and inclusive team environment.
  • Opportunities to participate in community and charity initiatives.
  • Global employee assistance programme supporting employee wellbeing.
  • Recognition through a global employee achievement platform.

Tech Stack

Categories

Site Reliability
Contact me