3 days ago
Bengaluru, India or Hyderābād, IndiaSenior
Responsibilities
- Engineer and operate a scalable monitoring and observability platform for hybrid cloud clients.
- Plan and execute the observability and monitoring tools roadmap.
- Define monitoring best practices for proactive alerting, anomaly detection, and performance analytics.
- Operate and optimize monitoring solutions across networks, distributed systems, and applications.
- Establish alerting thresholds based on Service Level Objectives and Service Level Agreements.
- Create monitoring audit standards for standard and custom monitors.
- Serve as an escalation point for monitoring-related incidents.
- Automate monitoring configurations and telemetry collection with Ansible, Terraform, and scripting.
Requirements
- At least 7 years of experience in observability or monitoring engineering operational roles.
- At least 7 years of hands-on experience with ITSM platforms and monitoring tools such as ServiceNow, BMC, Datadog, and Entuity.
- Strong proficiency in Python, Bash, and JavaScript for automation and scripting.
- Experience using Ansible, Terraform, or similar infrastructure-as-code tools for observability deployments.
- Strong analytical, problem-solving, communication, leadership, training, and cross-functional collaboration skills.
- Bachelor’s degree in a related field.
- Preferred: master’s degree in an information technology-related field.
- Preferred: proficiency with AWS, Azure, GCP, and Kubernetes deployment and monitoring.
- Preferred: advanced ITIL v3 or v4 certification or training.
- Preferred: experience integrating AI/ML into ITSM practices.
