Priceline

Site Reliability Engineer - Observability

Priceline
Apply
2 months ago
Mumbai, IndiaMid Level

Responsibilities

  • Support and evolve end-to-end observability solutions for collecting, shipping, storing, and querying OpenTelemetry metrics, logs, and traces across infrastructure, containers, and Kubernetes environments.
  • Administer and operate Splunk, New Relic, ClickHouse, Grafana, and Lightrun, including onboarding, access management, configuration, upgrades, reliability, performance, and SLAs.
  • Drive standardized logging, metrics, distributed tracing, schemas, instrumentation practices, and observability governance across services.
  • Partner with product, platform, and engineering teams to improve production visibility and support SLO-driven reliability practices.
  • Optimize telemetry pipelines for performance, data quality, scalability, and cost through sampling, filtering, and data lifecycle management.
  • Build dashboards and actionable alerts, configure alert workflows and incident-management integrations, and improve signal-to-noise ratio.
  • Lead complex incident investigations and postmortems, identify observability gaps, and drive improvements to reduce MTTR and MTTD.
  • Advance the observability platform toward intelligent and AI-enabled capabilities, including MCP-based solutions.

Requirements

  • Bachelor’s degree in Computer Science or equivalent practical experience.
  • At least 4 years of experience in observability, SRE, DevOps, or platform engineering roles supporting production systems.
  • Strong understanding of APM and SRE fundamentals, including metrics, events, logs, traces, latency analysis, error rates, service dependency mapping, SLIs, SLOs, alert tuning, and root cause analysis.
  • Hands-on experience administering at least one modern observability or APM platform such as Splunk, New Relic, or Grafana, including metrics, logs, distributed tracing, and platform configuration.
  • Experience with full-stack observability across infrastructure, applications, and browser monitoring layers and with operating platforms at scale.
  • Experience implementing or supporting OpenTelemetry instrumentation, improving telemetry quality, and reducing alert fatigue.
  • Familiarity with Kubernetes and cloud-native environments, including troubleshooting complex production issues in distributed systems.
  • Experience managing telemetry pipelines and agents such as collectors, forwarders, or sidecars, including onboarding services and resolving ingestion issues.
  • Working knowledge of Shell or Python scripting and CI/CD concepts.
  • Terraform experience for managing platform configurations and integrations is a plus.
  • Ability to collaborate with engineering teams, influence monitoring standards, evaluate observability cost and performance trade-offs, and lead incident investigations and postmortems.
  • Relevant certifications such as New Relic APM Professional, Reliability Engineer – Professional, Splunk Admin, or GCP Associate Cloud Engineer are a plus.
  • Demonstrated alignment with Priceline’s values of Customer, Innovation, Team, Accountability, and Trust.

Benefits

  • Hybrid work model with two days onsite, with the remaining days available remotely or in the office.
  • Medical, dental, vision, and mental health coverage.
  • Paid time off, holidays, a company-wide Priceline Pause reset week, and paid volunteer days.
  • Ability to work up to four weeks per year from anywhere, parental leave, dependent care and family support, Summer Fridays, stocked kitchens, and catered meals where available.
  • Retirement plans with company contributions, life and disability coverage, and tax-advantaged accounts.
  • Employee discounts on hotels and flights, VIP deals, and Big Deal Bucks travel credits.
  • Travel and partner discounts, tuition support, legal support, and pet benefits.
  • Employee Resource Groups, social events, recognition programs, and service awards.

Tech Stack

Categories

Site Reliability
Priceline

About Priceline

1,001-5,000 employees
Contact me