1 month ago
Mumbai, IndiaMid Level
Responsibilities
- Support and evolve observability solutions for collecting, shipping, storing, and querying OpenTelemetry metrics, logs, and traces across infrastructure, containers, and Kubernetes.
- Administer and operate Splunk, New Relic, ClickHouse, Grafana, and Lightrun, including onboarding, access management, configuration, upgrades, and reliability management.
- Drive standardized logging, metrics, distributed tracing, instrumentation, schemas, and conventions across services.
- Partner with product, platform, and engineering teams to improve production visibility and SLO-driven reliability practices.
- Optimize telemetry pipelines for performance, data quality, scalability, and cost through sampling, filtering, and data lifecycle management.
- Define observability governance standards and promote adoption through documentation, tooling, and enablement.
- Lead incident investigations and postmortems, identify observability gaps, and improve MTTR, MTTD, alert quality, and signal-to-noise ratio.
- Explore intelligent and AI-enabled observability capabilities, including MCP-based solutions.
Requirements
- Bachelor ’s degree in Computer Science or equivalent practical experience.
- At least 4 years of experience in observability, SRE, DevOps, or platform engineering roles supporting production systems.
- Strong understanding of APM, SRE fundamentals, MELT, latency analysis, error monitoring, dependency mapping, SLIs, SLOs, alert tuning, and root-cause analysis in distributed systems.
- Hands-on administration experience with at least one modern observability or APM platform such as Splunk, New Relic, or Grafana.
- Experience supporting observability across infrastructure, application, and browser monitoring layers at scale.
- Experience building dashboards and actionable alerts and integrating alert workflows with incident-management tools such as PagerDuty.
- Experience implementing or supporting OpenTelemetry instrumentation and improving telemetry quality.
- Familiarity with Kubernetes and cloud-native environments, including troubleshooting complex production issues.
- Experience managing telemetry pipelines and agents such as collectors, forwarders, and sidecars.
- Working knowledge of Shell or Python scripting and CI/CD concepts.
- Experience with infrastructure-as-code tools such as Terraform is preferred.
- Experience leading or contributing to incident investigations and postmortems, with the ability to influence engineering teams and balance observability depth, performance, and cost.
- Relevant certifications such as New Relic APM Professional, Reliability Engineer – Professional, Splunk Admin, or GCP Associate Cloud Engineer are preferred.
- Demonstrated alignment with Priceline’s values and high standards of ethics, honesty, transparency, and compliance.
Benefits
- Hybrid work model with two days onsite and remaining days remote or in the office.
- Medical, dental, vision, and mental health coverage.
- Paid time off, holidays, a company-wide Priceline Pause reset week, and paid volunteer days.
- Ability to work up to four weeks per year from anywhere, plus parental leave, dependent care, family support resources, Summer Fridays, stocked kitchens, and catered meals where available.
- Retirement plans with company contributions, life and disability coverage, and tax-advantaged accounts.
- Employee discounts on hotels and flights, VIP deals, and Big Deal Bucks travel credits.
- Additional travel and partner discounts, tuition support, legal support, and pet benefits.
- Employee Resource Groups, social events, recognition programs, and service awards.
Tech Stack
Categories
Site Reliability
