Priceline

Site Reliability Engineer, Observability

Priceline
Apply
1 month ago
Mumbai, IndiaMid Level

Responsibilities

  • Support and evolve observability solutions for collecting, shipping, storing, and querying OpenTelemetry metrics, logs, and traces across infrastructure, containers, and Kubernetes.
  • Administer and operate Splunk, New Relic, ClickHouse, Grafana, and Lightrun, including onboarding, access management, configuration, upgrades, and reliability management.
  • Drive standardized logging, metrics, distributed tracing, instrumentation, schemas, and conventions across services.
  • Partner with product, platform, and engineering teams to improve production visibility and SLO-driven reliability practices.
  • Optimize telemetry pipelines for performance, data quality, scalability, and cost through sampling, filtering, and data lifecycle management.
  • Define observability governance standards and promote adoption through documentation, tooling, and enablement.
  • Lead incident investigations and postmortems, identify observability gaps, and improve MTTR, MTTD, alert quality, and signal-to-noise ratio.
  • Explore intelligent and AI-enabled observability capabilities, including MCP-based solutions.

Requirements

  • Bachelor’s degree in Computer Science or equivalent practical experience.
  • At least 4 years of experience in observability, SRE, DevOps, or platform engineering roles supporting production systems.
  • Strong understanding of APM, SRE fundamentals, MELT, latency analysis, error monitoring, dependency mapping, SLIs, SLOs, alert tuning, and root-cause analysis in distributed systems.
  • Hands-on administration experience with at least one modern observability or APM platform such as Splunk, New Relic, or Grafana.
  • Experience supporting observability across infrastructure, application, and browser monitoring layers at scale.
  • Experience building dashboards and actionable alerts and integrating alert workflows with incident-management tools such as PagerDuty.
  • Experience implementing or supporting OpenTelemetry instrumentation and improving telemetry quality.
  • Familiarity with Kubernetes and cloud-native environments, including troubleshooting complex production issues.
  • Experience managing telemetry pipelines and agents such as collectors, forwarders, and sidecars.
  • Working knowledge of Shell or Python scripting and CI/CD concepts.
  • Experience with infrastructure-as-code tools such as Terraform is preferred.
  • Experience leading or contributing to incident investigations and postmortems, with the ability to influence engineering teams and balance observability depth, performance, and cost.
  • Relevant certifications such as New Relic APM Professional, Reliability Engineer – Professional, Splunk Admin, or GCP Associate Cloud Engineer are preferred.
  • Demonstrated alignment with Priceline’s values and high standards of ethics, honesty, transparency, and compliance.

Benefits

  • Hybrid work model with two days onsite and remaining days remote or in the office.
  • Medical, dental, vision, and mental health coverage.
  • Paid time off, holidays, a company-wide Priceline Pause reset week, and paid volunteer days.
  • Ability to work up to four weeks per year from anywhere, plus parental leave, dependent care, family support resources, Summer Fridays, stocked kitchens, and catered meals where available.
  • Retirement plans with company contributions, life and disability coverage, and tax-advantaged accounts.
  • Employee discounts on hotels and flights, VIP deals, and Big Deal Bucks travel credits.
  • Additional travel and partner discounts, tuition support, legal support, and pet benefits.
  • Employee Resource Groups, social events, recognition programs, and service awards.

Tech Stack

Categories

Site Reliability
Priceline

About Priceline

1,001-5,000 employees
Contact me