3 months ago
Bengaluru, IndiaStaff+
Responsibilities
- Drive the technical vision and architecture for Postman’s observability platform.
- Design and build scalable solutions for metrics, logging, tracing, alerting, and operational analytics.
- Investigate complex production issues, identify root causes, and drive long-term corrective actions.
- Improve service reliability, availability, performance, and operational maturity in partnership with engineering teams.
- Establish observability standards, best practices, and instrumentation frameworks across the company.
- Build tooling and automation for faster incident detection, diagnosis, and resolution.
- Use telemetry data to identify performance bottlenecks, capacity risks, and reliability gaps.
- Lead cross-functional initiatives focused on platform health, operational excellence, and engineering productivity.
- Mentor senior engineers and raise the technical bar across the organization.
Requirements
- 10+ years of software engineering experience with significant exposure to distributed systems and cloud-native architectures.
- Strong expertise in monitoring, logging, distributed tracing, telemetry pipelines, and incident management.
- Experience operating large-scale production systems with a focus on reliability, scalability, and performance.
- Deep understanding of system debugging, root-cause analysis, performance optimization, and production operations.
- Strong programming experience in one or more languages such as Go, Java, Python, or Node.js.
- Experience with observability technologies such as OpenTelemetry, Prometheus, Grafana, Elasticsearch, Datadog, New Relic, Splunk, or Honeycomb.
- Ability to influence technical direction across teams without direct authority.
- Strong communication skills and a data-driven approach to problem-solving.
- Preferred experience building internal developer platforms or observability platforms at scale.
- Preferred exposure to AIOps, intelligent alerting, anomaly detection, or AI-powered operational tooling.
- Preferred experience driving reliability initiatives across multiple engineering organizations.
- Preferred background in SRE, Platform Engineering, Infrastructure Engineering, or Developer Productivity.
Benefits
- Flexible schedule
- Full medical coverage
- Flexible paid time off (PTO)
- Wellness reimbursement
- Monthly lunch stipend
- Wellness programs
- Team-building events
- Donation-matching program
- In-office work five days per week for roles based in San Francisco Bay Area, Boston, Austin, New York City, Tokyo, and London
- Bangalore roles currently require three days per week in the office and are expected to transition to five days per week by year-end
Categories
DevOpsSite Reliability
About Postman
Postman builds a SaaS platform for designing, testing, documenting, and managing APIs used by software teams across development and operations. The company sells free and paid plans, including enterprise features for collaboration, security, and governance, and offers desktop and cloud tools. Founded in 2014 and headquartered in San Francisco, Postman is privately held and reports adoption by over 40 million developers and 500,000 organizations worldwide.
