
Lead Engineer - Reliability Engineering
StoneX Group Inc.2 hours ago
Bengaluru, IndiaStaff+
Responsibilities
- Define and drive reliability engineering standards, practices, maturity models, service tiering, and adoption metrics across platforms and services.
- Partner with engineering, platform, infrastructure, and product teams to improve reliability, resilience, operability, supportability, and service ownership.
- Establish SLOs, SLIs, error budgets, alert-quality standards, toil-reduction practices, production-readiness reviews, and service ownership expectations.
- Improve change safety, release confidence, capacity planning, resilience testing, disaster recovery, incident management, post-incident reviews, and on-call effectiveness.
- Drive end-to-end observability adoption, including metrics, logs, traces, service maps, dashboards, instrumentation, actionable alerting, and service health reporting.
- Define observability architecture and telemetry standards and embed reliability guardrails into CI/CD pipelines, internal developer platforms, templates, workflows, scorecards, and dashboards.
- Identify and address reliability risks, architectural weaknesses, service fragility, and third-party or provider dependencies.
- Define and track reliability metrics and operational KPIs that improve service outcomes over time.
- Contribute hands-on to technical design, implementation, and operational improvement while guiding technical direction.
- Mentor and coach engineers and help raise reliability maturity across the organization.
Requirements
- 7+ years of experience in SRE, production engineering, platform engineering, infrastructure engineering, or a closely related role.
- Track record of building, improving, or scaling reliability engineering or SRE practices in a large organization.
- Several years of hands-on experience supporting production systems at scale, including incident response, problem management, availability improvement, and operational excellence.
- Strong experience defining and implementing SLOs, SLIs, error budgets, service health models, and reliability-focused engineering practices.
- Strong experience with observability platforms such as Datadog, including metrics, logs, tracing, alerting, dashboards, and service-level reporting.
- Experience driving end-to-end observability adoption across application teams, including instrumentation, telemetry standards, dashboards, alerting, and service-level reporting.
- Experience automating operational processes with Terraform, scripting languages, CI/CD pipelines, and cloud-native platforms.
- Experience with Kubernetes, Linux, Git, and modern cloud or platform infrastructure.
- Strong systems thinking and the ability to balance reliability, latency, engineering velocity, risk, and cost.
- Strong communication, collaboration, influencing, mentoring, and cross-functional leadership skills.
- Bachelor’s degree in computer science, engineering, or a related field, or equivalent practical experience.
- Relevant certifications are a plus, along with a commitment to continual professional and technical development.
Benefits
- Four days’ work from the office.
- Occasional travel for team collaboration meetings and conferences.
Tech Stack
Categories
Site Reliability