Razorpay

Site Reliability Engineer

Razorpay
Apply
4 hours ago
Bengaluru, IndiaSenior

Responsibilities

  • Build and operate reliable, fault-tolerant infrastructure for Razorpay’s payment, payout, reconciliation, and other production systems.
  • Develop expertise in rate limiting, sharding, traffic management, and large-scale distributed-system reliability.
  • Lead capacity planning, failure-mode analysis, architectural improvements, and production reliability projects.
  • Design automation tooling and self-healing infrastructure to reduce manual toil and improve resilience.
  • Build AI- and ML-driven anomaly detection, predictive alerting, noise reduction, and observability systems.
  • Use LLM-based tooling for incident triage, root-cause analysis, incident summarization, runbook generation, and post-incident learning.
  • Build and maintain reliable infrastructure for AI-native products, ML models, and inference pipelines.
  • Provide on-call support, participate in incident response and blameless postmortems, and drive remediation.
  • Evaluate emerging AIOps tools, observability platforms, and reliability patterns.
  • Mentor engineers and promote SRE practices including SLOs, error budgets, toil reduction, and data-driven reliability decisions.

Requirements

  • Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
  • At least 5 years of experience in software engineering, systems engineering, or site reliability engineering, including 3 years focused on site reliability engineering.
  • At least 3 years of experience in software design and architecture, including distributed systems and backend services.
  • Production-quality programming and automation experience in at least one language such as Go, Python, or Java.
  • Experience using AI-assisted development and operations tooling, including LLM-based copilots for coding, debugging, and incident response.
  • Preferred: Master’s degree in Computer Science or a related technical field.
  • Preferred: 5+ years of large-scale distributed-systems experience, ideally in payments, fintech, e-commerce, or cloud infrastructure.
  • Preferred: Experience with AIOps, ML-based anomaly detection, intelligent alerting, LLM-powered operations tooling, and infrastructure for ML/AI workloads.
  • Preferred: Experience with chaos engineering, fault injection, resilience testing, cloud-native infrastructure, Kubernetes, service mesh, cloud platforms, and infrastructure-as-code.
  • Preferred: Experience troubleshooting complex distributed systems under production pressure and influencing stakeholders across engineering and product.

Categories

Site Reliability
Razorpay

About Razorpay

1,001-5,000 employees

Razorpay builds a full-stack payments and banking platform for businesses in India and Southeast Asia, offering payment gateway, payouts, subscriptions, payroll, neobanking (RazorpayX), and working capital (Razorpay Capital). Founded in 2014 and headquartered in Bangalore, it earns via transaction fees and financial services, and processes over $180 billion in annualized payment volume for customers across e-commerce, telecom, and consumer internet.

Contact me