2 months ago
Mumbai, IndiaMid Level
Responsibilities
- Support and evolve end-to-end observability solutions for collecting, shipping, storing, and querying OpenTelemetry metrics, logs, and traces across infrastructure, containers, and Kubernetes environments.
- Administer and operate Splunk, New Relic, ClickHouse, Grafana, and Lightrun, including onboarding, access management, configuration, upgrades, reliability, performance, and SLAs.
- Drive standardized logging, metrics, distributed tracing, schemas, instrumentation practices, and observability governance across services.
- Partner with product, platform, and engineering teams to improve production visibility and support SLO-driven reliability practices.
- Optimize telemetry pipelines for performance, data quality, scalability, and cost through sampling, filtering, and data lifecycle management.
- Build dashboards and actionable alerts, configure alert workflows and incident-management integrations, and improve signal-to-noise ratio.
- Lead complex incident investigations and postmortems, identify observability gaps, and drive improvements to reduce MTTR and MTTD.
- Advance the observability platform toward intelligent and AI-enabled capabilities, including MCP-based solutions.
Requirements
- Bachelor’s degree in Computer Science or equivalent practical experience.
- At least 4 years of experience in observability, SRE, DevOps, or platform engineering roles supporting production systems.
- Strong understanding of APM and SRE fundamentals, including metrics, events, logs, traces, latency analysis, error rates, service dependency mapping, SLIs, SLOs, alert tuning, and root cause analysis.
- Hands-on experience administering at least one modern observability or APM platform such as Splunk, New Relic, or Grafana, including metrics, logs, distributed tracing, and platform configuration.
- Experience with full-stack observability across infrastructure, applications, and browser monitoring layers and with operating platforms at scale.
- Experience implementing or supporting OpenTelemetry instrumentation, improving telemetry quality, and reducing alert fatigue.
- Familiarity with Kubernetes and cloud-native environments, including troubleshooting complex production issues in distributed systems.
- Experience managing telemetry pipelines and agents such as collectors, forwarders, or sidecars, including onboarding services and resolving ingestion issues.
- Working knowledge of Shell or Python scripting and CI/CD concepts.
- Terraform experience for managing platform configurations and integrations is a plus.
- Ability to collaborate with engineering teams, influence monitoring standards, evaluate observability cost and performance trade-offs, and lead incident investigations and postmortems.
- Relevant certifications such as New Relic APM Professional, Reliability Engineer – Professional, Splunk Admin, or GCP Associate Cloud Engineer are a plus.
- Demonstrated alignment with Priceline’s values of Customer, Innovation, Team, Accountability, and Trust.
Benefits
- Hybrid work model with two days onsite, with the remaining days available remotely or in the office.
- Medical, dental, vision, and mental health coverage.
- Paid time off, holidays, a company-wide Priceline Pause reset week, and paid volunteer days.
- Ability to work up to four weeks per year from anywhere, parental leave, dependent care and family support, Summer Fridays, stocked kitchens, and catered meals where available.
- Retirement plans with company contributions, life and disability coverage, and tax-advantaged accounts.
- Employee discounts on hotels and flights, VIP deals, and Big Deal Bucks travel credits.
- Travel and partner discounts, tuition support, legal support, and pet benefits.
- Employee Resource Groups, social events, recognition programs, and service awards.
Tech Stack
Categories
Site Reliability
