3 months ago
Bengaluru, IndiaStaff+
Responsibilities
- Design and build platforms, tools, and frameworks that improve system reliability, scalability, and performance.
- Define and implement SRE practices, including SLIs, SLOs, error budgets, and reliability metrics.
- Lead incident response, root cause analysis, and long-term remediation efforts.
- Analyze system behavior, identify bottlenecks and saturation points, and improve resilience.
- Partner with engineering teams to embed reliability into the software development lifecycle.
- Evaluate emerging technologies and recommend tools that improve productivity, observability, and system robustness.
- Drive capacity planning, performance tuning, and cost optimization.
- Collaborate with cross-functional teams to identify gaps, prioritize improvements, and resolve production issues.
- Provide technical leadership and mentorship across the engineering organization.
- Influence senior leadership with system health insights, metrics, and operational recommendations.
Requirements
- Bachelor’s or master’s degree in Computer Science, Engineering, or a related technical field.
- 10+ years of software engineering experience with a strong focus on backend systems and distributed architecture.
- Extensive experience building and operating Java-based systems using Spring Boot.
- Strong understanding of distributed systems, fault tolerance, eventual consistency, and scalability.
- Proven experience with AWS, Azure, or GCP and cloud-native architectures.
- Expertise with observability tools such as Prometheus, Grafana, and ELK or similar tools.
- Experience defining and managing SLIs, SLOs, and error budgets.
- Strong knowledge of automation and infrastructure as code.
- Hands-on experience with incident management, root cause analysis, and postmortems.
- Excellent analytical, debugging, problem-solving, communication, collaboration, and leadership abilities.
Benefits
- Office-first culture encouraging three days per week in the office for most roles, with role-specific flexibility confirmed with the recruiter.
- Comprehensive healthcare coverage, flexible paid time off, equity RSUs, annual performance bonus opportunities, retirement account support, and 14+ weeks of paid parental leave.
- Career development opportunities and company-paid privacy certification exam fees.
- Benefits vary by country.
Tech Stack
Categories
BackendSite Reliability
About OneTrust
OneTrust, the AI-Ready Governance Platform™, enables innovation through the responsible use of data and AI. Trusted by over half of the Fortune 500, we help businesses govern well and move fast, turning responsible data use into a catalyst for growth. To learn more, visit www.onetrust.com