4 months ago
Hyderābād, IndiaSenior
Responsibilities
- Lead client discovery sessions, assessments, and workshops covering observability, telemetry, reliability, and operational maturity.
- Define observability architectures and roadmaps aligned with cloud, platform engineering, SRE, and AIOps initiatives.
- Design and guide implementations across metrics, logs, traces, dashboards, alerting, and incident workflows.
- Build or refine dashboards, alerts, service views, operational integrations, and telemetry governance practices.
- Improve alert quality, incident triage, escalation paths, runbooks, post-incident feedback loops, and reliability measurement.
- Advise clients on telemetry data quality, retention, cardinality, access, sampling, and cost optimization.
- Partner with engineering, platform, operations, and leadership stakeholders to align observability investments with business priorities.
- Provide hands-on technical leadership through configuration guidance, implementation oversight, design validation, troubleshooting, and quality review.
- Support integrations across observability, ITSM, incident management, collaboration, workflow, and AI-assisted operations platforms.
- Contribute to implementation planning, solution governance, documentation, enablement, operational handoff, reusable accelerators, and practice best practices.
- Mentor delivery teams and lead client-facing discussions across technical and executive audiences.
Requirements
- 6+ years of experience in consulting, engineering, SRE, platform engineering, or operations with strong observability responsibility.
- Hands-on experience with metrics, logs, traces, alerting, service health, incident response, OpenTelemetry, telemetry pipelines, instrumentation, and collection architecture.
- Working knowledge of open-source observability tools such as Prometheus, Grafana, Loki, Tempo, Mimir, Jaeger, or Elastic.
- Experience with one or more enterprise observability platforms such as Datadog, Dynatrace, Splunk, New Relic, Elastic, LogicMonitor, Honeycomb, or Chronosphere.
- Strong understanding of Kubernetes, containers, cloud-native architectures, distributed systems, and modern application architectures.
- Experience with Terraform, Helm, CI/CD pipelines, Ansible, or related automation and platform tooling.
- Understanding of SRE, DevOps, ITSM, service health models, SLIs, SLOs, alerting strategies, reliability measurement, and business outcomes.
- Experience with error budgets, burn-rate alerting, production readiness reviews, reliability governance, telemetry cost optimization, sampling, retention, tagging, cardinality management, or observability governance.
- Experience integrating observability with ServiceNow, incident management platforms, CMDBs, or collaboration tools.
- Familiarity with platform engineering, internal developer platforms, service catalogs, golden paths, self-service observability, AIOps, anomaly detection, event correlation, automated remediation, or MCP-enabled integrations.
- Familiarity with microservices, service mesh technologies, end-user monitoring, or digital experience monitoring.
- Generalist coding or scripting experience in Python, Java, Go, JavaScript, or .NET.
- Ability to lead client-facing conversations, structure ambiguous problems, communicate tradeoffs, and translate strategy into executable workstreams.
- Strong written and verbal communication skills with technical and executive audiences, along with a collaborative continuous-learning approach.
Benefits
- Comprehensive health insurance coverage in India, with options to extend coverage to dependents.
- Paid time off, company holidays, and additional leave benefits according to policy.
- Flexible work arrangements supporting work-life balance.
- Learning and development opportunities, including sponsored certifications and credentials.
- Employee wellness initiatives focused on physical and mental well-being.
- Retirement and statutory benefits in line with India regulations.
- Inclusive, collaborative, people-first culture with opportunities for ownership and practice growth.
Tech Stack
Categories
Solutions Engineering
